top of page

Microsoft launches Decision-1 model with 35× faster inference than GPT-6 Sol for AI agent decisions

2 minutes ago
10 min read
Microsoft Decision-1 model launch, AI agent decisions and 35x faster inference - Data Studios

Microsoft has introduced Microsoft-Decision-1, a specialized artificial intelligence model designed to make structured decisions at substantially lower latency and cost than conventional large language models.


Announced on October 9, 2026, the model is available through Microsoft Foundry and OpenRouter, targeting developers building autonomous AI agents, classification systems, intelligent routing mechanisms, automated evaluation pipelines, and enterprise workflows.


According to Microsoft's internal evaluations, Decision-1 achieved the highest overall accuracy among the models tested across 36 benchmarks covering nearly 150,000 questions, while delivering median decision latency approximately 35 times lower than GPT-6 Sol under the company's testing conditions.


The model also introduces an unusual pricing structure: $0.042 per million input tokens, with no charge for output tokens. This positions Decision-1 as a potentially economical component for applications executing large numbers of recurring classification and decision-making operations.


Rather than competing directly with general-purpose models in conversation, coding, or complex reasoning, Microsoft-Decision-1 addresses a different workload: selecting among predefined alternatives and returning probability scores that applications can interpret and act upon.


Its introduction reflects a broader architectural shift toward separating the reasoning capabilities of large AI models from the narrower, repetitive decisions required to coordinate their actions.


··········


MICROSOFT-DECISION-1 USES A 9B-PARAMETER FOUNDATION FOR STRUCTURED DECISION SCORING.


The model is built on Alibaba's Qwen3.5-9B architecture, which Microsoft has post-trained to evaluate predefined options through a single inference pass.


Unlike conventional language models that generate text token by token, Decision-1 processes an input situation together with a closed set of available answers.


It then returns a probability score for each answer, allowing the surrounding application to determine which alternative should be selected.


Supported formats include binary decisions, multiple-choice classification, numerical ratings, relevance evaluation, and rubric-based assessments of AI-generated responses or proposed agent actions.


For example, a customer-support system might classify a request into billing, technical assistance, account security, or human escalation. A software agent could use the same mechanism to determine whether an attempted operation satisfies predefined safety requirements.


The model's output is intended to be consumed programmatically rather than displayed as a conversational response.


Microsoft's implementation incorporates probability calibration, meaning that the returned scores are designed to reflect the estimated likelihood that an answer is correct. This allows applications to establish confidence thresholds rather than treating every prediction as equally reliable.


The model currently uses a 32,768-token context window, accepts text-only input, and produces structured JSON output.


........


Technical characteristic

Microsoft-Decision-1

Developer

Microsoft

Release date

October 9, 2026

Foundation model

Qwen3.5-9B

Foundation model size

9 billion parameters

Training approach

Microsoft post-training

Context window

32,768 tokens

Supported input

Text

Output

Structured probability scores

Decision formats

Yes/no, multiple-choice, ratings, rubrics

Inference approach

Single-pass decision scoring

Availability

Microsoft Foundry and OpenRouter

Model distribution

Hosted proprietary service

Input pricing

$0.042 per million tokens

Output pricing

Free


........


Although Qwen3.5-9B is an open-weight foundation model, Microsoft's post-trained Decision-1 is distributed as a hosted service rather than as downloadable model weights.


Microsoft also plans to develop subsequent versions using other foundations, including models from Microsoft AI and OpenAI.


The underlying model architecture may therefore change over time while the service continues to expose the same decision-scoring functionality.


··········


HOW MICROSOFT-DECISION-1 CONTROLS AI AGENTS AND AUTOMATED WORKFLOWS.


Decision-1 is designed to operate between reasoning, execution, and verification stages, where an AI application must choose a specific action without generating another extensive natural-language response.


Modern agentic systems frequently require dozens of decisions during a single task.


An agent retrieving information might need to select a search tool, evaluate document relevance, determine whether sufficient evidence has been collected, verify the reliability of a generated answer, and decide whether another operation is necessary.


Using a large reasoning model for every intermediate decision can substantially increase execution time and inference expenditure.


Decision-1 offers a specialized alternative for steps where the available options are already known.


Consider a software agent preparing to modify an enterprise database.


Before executing the operation, the application could submit the proposed action, relevant authorization information, and organizational policy requirements to Decision-1.


The model might evaluate four predefined outcomes:


  • Continue: The operation satisfies the supplied requirements.

  • Retry: The proposed operation should be revised before execution.

  • Escalate: Additional human authorization or investigation is necessary.

  • Block: The operation should not proceed.


The resulting probabilities can inform a decision policy implemented by the application.


For instance, a sufficiently high confidence score for an authorized operation might permit execution, while an ambiguous result would require human review.


These are illustrative application-level choices rather than a claim that Decision-1 independently enforces access controls.


The model does not automatically execute tools, establish user permissions, or guarantee that a proposed action is safe. Those responsibilities remain with the host application and its security infrastructure.


This distinction is important because a model-generated confidence score is not equivalent to formal authorization.


For higher-risk operations, deterministic permission checks and independent safeguards remain necessary regardless of the model's assessment.


··········


MICROSOFT REPORTS 35× LOWER LATENCY THAN GPT-6 SOL ACROSS STRUCTURED DECISION TASKS.


The headline performance advantage comes from Microsoft's own benchmarking environment and concerns decision-scoring latency, rather than general-purpose reasoning or text generation.


Microsoft evaluated Decision-1 against language models and specialized decision systems using 36 public and private benchmarks spanning nearly 150,000 questions.


The evaluated workloads included routing, ranking, multilingual classification, long-context decisions, reasoning-related assessments, safety evaluation, and tasks outside the training distribution.


Microsoft reports that Decision-1 achieved the highest overall accuracy among the models included in these evaluations.


The company also measured substantial improvements in inference speed.


........


Performance indicator

Microsoft-reported result

Evaluation benchmarks

36

Questions evaluated

Nearly 150,000

Overall accuracy ranking

Highest among tested models

Median latency vs GPT-6 Sol

Approximately 35× faster

Latency vs H2O-Lightning-4B v1.1

Approximately 2.5× faster

Average decision changes under input perturbations

1.3%

Tested perturbation types

8

Safety-evaluation requests

5,250

Safety benchmarks

11


........


The speed comparison refers to median latency, commonly represented as P50. This measures the response time at which half the evaluated requests complete faster and half complete more slowly.


It does not establish that Decision-1 is 35 times faster than GPT-6 Sol on arbitrary tasks.


A general-purpose reasoning model may perform analysis, generate explanations, interact with tools, and solve open-ended problems that Decision-1 is not designed to handle.


The comparison is therefore most relevant when both models are used to produce a decision from a predefined set of alternatives.


Microsoft also examined robustness under equivalent input transformations.


The company introduced eight types of perturbations, including changes to option order, paraphrased descriptions, and formatting differences.


Across these tests, Decision-1 changed its selected answer on an average of 1.3% of perturbations. Microsoft reported no decision changes when answer options were reordered or when their descriptions were paraphrased in the tested cases.


This behavior is valuable in production systems because semantically equivalent requests can otherwise generate inconsistent classifications.


However, robustness against formatting changes does not establish resistance to adversarial manipulation or guarantee correct predictions.


Microsoft also evaluated the model on 5,250 requests across 11 safety benchmarks addressing harmful content, jailbreak attempts, and prompt injection.


The company reports favorable safety performance, although those results should be distinguished from independently demonstrated security guarantees.


More broadly, the published benchmark comparisons remain vendor-conducted evaluations. Independent testing under standardized conditions will be necessary to establish how Decision-1 performs across different providers, application environments, and real-world datasets.


··········


PRICING STARTS AT $0.042 PER MILLION INPUT TOKENS WITH FREE OUTPUT.


Microsoft-Decision-1 uses input-token billing rather than conventional two-sided input and output pricing, reflecting its specialization in compact structured responses.


Microsoft lists the model at $0.042 per million input tokens, while output tokens are free.


For organizations processing large numbers of short requests, this creates an unusually low direct inference cost.


The total expense depends primarily on the amount of input context supplied to the model.


An application passing short instructions, a small number of answer options, and limited supporting evidence can operate at very low marginal cost.


Conversely, workloads involving lengthy documents, complex policy descriptions, or substantial contextual information will consume more input tokens.


The following estimates illustrate the model's published token pricing.


........


Monthly decisions

Average input tokens

Estimated inference cost

10,000

500

$0.21

100,000

500

$2.10

1,000,000

500

$21.00

1,000,000

1,000

$42.00

10,000,000

500

$210.00


........


These calculations assume that every request uses the specified number of billable input tokens and that the published price remains unchanged.


They exclude ancillary infrastructure, networking, application orchestration, monitoring, data storage, and any additional inference performed by other models.


They also exclude the engineering cost of establishing reliable decision criteria and handling uncertain predictions.


Free output tokens do not make an entire agent workflow free. Decision-1 may be inexpensive individually, but the surrounding system can still require substantial computation and external services.


An additional economic consideration is the cost per correct decision.


A lower-priced model is not necessarily more economical if it generates enough incorrect classifications to require additional model calls, human intervention, or operational corrections.


Production evaluations should consequently measure inference expenditure alongside decision accuracy, escalation frequency, and the cost of downstream errors.


··········


MICROSOFT IS ALREADY TESTING DECISION-1 IN XBOX RESEARCH, COPILOT, INCIDENT RESPONSE, AND SCIENTIFIC DISCOVERY.


Microsoft has disclosed several internal experiments demonstrating how specialized decision scoring can replace repeated language-model evaluations in existing workflows.


One application involves Xbox Research, which processes large volumes of user feedback to identify recurring themes, customer preferences, and product-related issues.


Researchers used Decision-1 to classify more than 10,000 pieces of open-ended feedback and reviews collected from surveys, Steam, and social platforms.


The model assigned the content to categories established by researchers, allowing large datasets to be organized without individually generating a natural-language explanation for each entry.


Microsoft reports that Decision-1 delivered classification quality competitive with GPT-6 Sol while operating more than 14 times faster and at approximately 200 times lower cost in this application.


These figures concern the specific Xbox Research workload and should not be interpreted as universal performance ratios.


Another application involves the Copilot team, which evaluates the quality of conversational and agent-generated outputs.


Microsoft reports that Decision-1 achieved quality competitive with GPT-5.6 Luna while operating approximately 100 times faster in the team's testing.


The model can assess whether a response satisfies a predefined rubric, allowing the system to accept suitable outputs or route uncertain results toward additional processing.


Microsoft has also tested Decision-1 in incident-response workflows.


Engineers investigating operational incidents frequently retrieve information from logs, support tickets, internal messages, and other enterprise data sources.


Decision-1 can help evaluate which information is relevant to the incident, allowing the retrieval system to prioritize useful material without invoking a general-purpose model for every classification.


A fourth application involves Microsoft Discovery, the company's scientific research platform.


The platform uses adaptive replanning, where agents evaluate experimental results, reassess previous actions, and determine whether another iteration is required.


Microsoft reports that Decision-1 produced scoring results 46 times more consistent than an LLM-based approach while evaluating those decisions three times faster.


In the tested adaptive-replanning workflow, this contributed to an almost fourfold improvement in execution speed.


These internal cases demonstrate the potential value of using specialized evaluation models inside larger AI systems.


They do not establish equivalent benefits for every enterprise workload, particularly where decisions require knowledge beyond the supplied evidence or substantial open-ended reasoning.


··········


LIMITATIONS INCLUDE CLOSED-ENDED INPUTS, CALIBRATION RISKS, AND THE ABSENCE OF EXPLANATORY REASONING.


Decision-1 is intended to select among predefined options, and its architecture imposes important restrictions that developers must consider before integrating it into production applications.


The most fundamental limitation is the requirement for a closed decision space.


The model is not designed for general conversation, document summarization, translation, open-ended question answering, or the generation of detailed explanations.


It also cannot independently retrieve missing information unless the surrounding application supplies that information through an external workflow.


This makes input preparation critical.


If an application presents incomplete evidence, poorly defined categories, or an inadequate set of possible answers, the model may produce a confident classification without resolving the underlying ambiguity.


Developers can introduce an explicit abstention option, such as an insufficient-evidence category, but this must be incorporated into the decision design.


Probability calibration presents another challenge.


A calibrated model attempts to align confidence scores with actual prediction accuracy. For example, predictions assigned approximately 90% confidence should be correct about 90% of the time across representative observations.


That relationship can deteriorate when the production data differs substantially from the evaluation distribution.


Changes in user behavior, input formatting, business policies, language distribution, and underlying model versions can all affect reliability.


Consequently, applications using Decision-1 should evaluate calibration on their own datasets and monitor performance over time.


The absence of explanatory output creates additional constraints for auditing and error investigation.


A probability score can indicate which alternative the model prefers, but it does not provide a detailed justification establishing why that alternative was selected.


For regulated or consequential processes, organizations may therefore need separate systems for evidence collection, rule validation, human review, and decision documentation.


Microsoft explicitly states that Decision-1 should not serve as the sole automated decision-maker in consequential decisions involving employment, credit, insurance, housing, healthcare, education, legal rights, or similar domains.


It is also not designed for surveillance, profiling, or tracking individuals.


These restrictions limit how the model should be deployed in sensitive settings even when its classification accuracy appears strong.


··········


DECISION MODELS COULD CHANGE HOW ENTERPRISES DESIGN AND OPERATE MULTI-AGENT SYSTEMS.


Microsoft-Decision-1 introduces a specialized inference layer that can reduce reliance on general-purpose language models for repetitive control and evaluation tasks.


An enterprise AI system may require sophisticated reasoning only during selected stages of an operation.


Other stages involve narrower questions: selecting a tool, checking whether retrieved evidence satisfies a condition, evaluating a proposed response, or determining whether a workflow should continue.


Handling every stage with the same large model can create unnecessary latency and expenditure.


A specialized decision model makes it possible to divide these responsibilities.


General-purpose models can handle planning, synthesis, coding, and complex analysis, while Decision-1 evaluates intermediate results and selects among permitted alternatives.


Traditional deterministic software remains preferable when a decision can be expressed through explicit, reliable rules without needing semantic interpretation.


Decision-1 occupies the space between these two approaches: situations requiring contextual judgment but producing a bounded, machine-readable outcome.


The architecture also introduces additional engineering requirements.


Each model invocation creates another dependency, while incorrect decisions can propagate into later workflow stages.


Applications must therefore distinguish reversible actions from irreversible operations, establish appropriate confidence thresholds, and ensure that safety-critical checks cannot be bypassed by model output alone.


The value of the model will depend on the balance between faster inference, classification quality, operational complexity, and downstream reliability.


For high-volume workloads, its pricing suggests substantial potential savings compared with repeatedly invoking larger models.


For complex or ambiguous tasks, however, the reduced computational cost may be offset by additional validation, escalation, or reasoning requirements.


Microsoft plans to update Decision-1 over time and incorporate additional training and evaluation data. The company has also indicated that future versions may use different underlying model families.


The launch therefore represents an initial commercial implementation of a broader approach to AI infrastructure rather than a fixed final architecture.


As autonomous systems become more complex, the economics of their intermediate decisions will increasingly influence deployment costs and execution speed.


Microsoft-Decision-1 demonstrates how a narrowly optimized model can provide useful decision intelligence at a fraction of the cost of general-purpose inference, provided that applications retain responsibility for validation, authorization, and the consequences of automated actions.


··········


FOLLOW US FOR MORE.


DATA STUDIOS


datastudios.org

bottom of page