After large language models (LLMs) and AI agents, a new category has emerged in the AI industry—decision modelsOn October 9, Microsoft officially launched Microsoft-Decision-1a lightweight model designed specifically for “fast decision-making,” now available on Microsoft Foundry and OpenRouter. Its positioning is clear: not to replace GPT-6 for writing articles or code, but to assist agents and business systems in completingrouting, classification, priority judgment, and result verification—high-frequency, low-cost decision tasks. Its official core selling point is striking: top accuracy across 36 blind-test benchmarks, while also being35× faster than GPT-6 Soland costing only $0.042 per million input tokens—with output completely free.
What Is a Decision Model? How Does It Differ from an LLM?
Understanding Microsoft-Decision-1 hinges on first clarifying the functional distinction between “decision models” and “large language models.” An LLM’s core strength lies ingenerating text and reasoning through complex problems—each invocation requires full autoregressive generation, resulting in high cost and latency; by contrast, a decision model isbuilt for structured output—it does not generate long-form text, but instead delivers judgments among a predefined set of options, producingstructured results(e.g., “yes/no,” a single option, or a score) that software can execute directly.
Microsoft officially positions it as a specialized model for four task categories:
- RoutingDetermines which model to route the request to
- ClassificationTags images, queries, and feedback
- Prioritization & VerificationScores results and determines whether human review is needed
- Workflow controlDetermines the next action for an Agent
This approach is not isolated. Just days ago, OpenAI launched Decisions APIa dedicated interface that separates “decision-making” from large language models; earlier, Amazon open-sourced Strands Decider 2Bvalidating the feasibility of running “decision models locally.” Microsoft’s entry effectively delivers a major endorsement for this emerging category.
Core capabilities of Microsoft-Decision-1
Technically, Microsoft-Decision-1 is derived from post-training on Qwen3.5-9B following a “single forward pass, rapid scoring” approach. Given a fixed set of options, it outputs acalibrated probability scorefor each option, supporting “yes/no,” “multiple choice,” “scoring,” and “scoring AI responses or Agent actions against a rubric.” Microsoft plans to rebase it onto MAI and OpenAI models in the future.
In its technical blog, Microsoft highlights three capabilities that distinguish decision models from “bolting LLMs together.”
- SpeedDecision tasks are often sequential, step by step. Official benchmarks show P50 latency is approximately 35× faster than GPT-6 Sol. In other words, for the same judgment task, GPT-6 Sol takes 100 ms, while this model takes only about 3 ms.
- Generalization qualityAchieves the highest accuracy across 36 benchmarks and nearly 150,000 questions—and these benchmarks wereheld out(unseen) during training. This avoids “leaderboard-fitting” overfitting.
- RobustnessWhen the same request is perturbed in eight ways (rephrasing, shuffling answer options, adding noise), the decision flip rate averages just 1.3%; when answer options are shuffled or scrambled, the flip rate is zero.
There is another practical yet easily overlooked point:Probabilities themselves are part of the API—not just rankings. Applications can use them to decide whether to “execute immediately” or “escalate for human review.” Officially claimed: predictions with 90% confidence are correct nine times out of ten on representative samples.
Real-world performance: faster than GPT-6 Sol and 200× cheaper
Microsoft shared several internal real-world usage metrics—more persuasive than synthetic benchmarks alone:
- Xbox Research data labelingClassifies over 10,000 open-ended feedback entries and reviews (from surveys, Steam, Twitter/X) by topic. Quality matches GPT-6 Sol, but speed improves 14× and cost drops 200×.
- Copilot quality assuranceEvaluates chat and agent response quality—performance matches GPT-5.6 Luna, yet runs 100× faster.
- Incident responseRetrieves relevant knowledge from logs, tickets, and messages—faster and more accurate than LLMs.
- scientific discoveryIn adaptive replanning scenarios, scoring consistency is 46× higher than LLM-based solutions.
These scenarios share one key trait:No new content generation required—only judgmentThis is precisely where decision models shine—saving costs previously incurred by overusing LLMs for tasks that don’t require generative capacity
Microsoft-Decision-1 vs. other decision/generation models
| Model | position | Speed/cost highlights | This Site’s Rating |
|---|---|---|---|
| Microsoft-Decision-1 | Decision models: routing, classification, validation | 35× faster than GPT-6 Sol, with free output | 8.7 |
| Strands Decider 2B | Open-source decision model, runnable locally | 2B parameters, lightweight local deployment | — |
| GPT-6.1 Sol | General-purpose efficiency model: generation + reasoning | Used as a speed benchmark (35× slower) | 8.8 |
| GPT-6 Sol / Luna | OpenAI’s flagship efficiency duo | Versatile but relatively high decision cost | 9.0 |
| Claude Haiku 5.5 | Anthropic’s lightweight workhorse | 75% price reduction; also suitable for lightweight judgments | 8.7 |
The table above reveals a clear trend:Decision-oriented tasks are being decoupled from general-purpose LLMsand delegated to smaller, more specialized, and cheaper dedicated models. For developers, this means an additional component in Agent architectures that significantly reduces cost and improves speed; for end users, it’s typically imperceptible directly, yet indirectly results in faster response times and lower costs across AI applications
User experience and limitations
Advantages are prominentSimple integration (a single structured API call), supported on both Foundry and OpenRouter—the two leading platforms; pricing is highly favorable for high-frequency calls—$0.042 per million input tokens, free output, delivering order-of-magnitude cost advantages for large-scale classification or routing
Limitations must also be stated clearly:
- Itcannot generate contentnor is it adept at open-ended, long-horizon reasoning. Writing articles, coding, and complex multi-step reasoning still require LLMs like GPT-6Claude and similar models.
- It solves “select one from given options” problems,and the quality of option designdirectly impacts performance—initial prompt/option engineering effort is required.
- Currently based on Qwen3.5-9B, official performance data for Chinese and other non-English scenarios is not separately released; actual effectiveness is recommended to be validated first on your own data.
- As a new category, its ecosystem and best practices are still in early stages, with few ready-to-use reference cases available.
Overall Score
| 维度 | Score | evaluate |
|---|---|---|
| functional completeness | 9.0 / 10 | Covers yes/no, multiple-choice, scoring, and rubric-based evaluation—clear, well-defined use cases |
| 易用性 | 8.5 / 10 | Single API call, available on both Foundry and OpenRouter platforms—fast onboarding |
| Cost-effectiveness | 9.5 / 10 | Output is free; input is extremely inexpensive—massive cost advantage for high-frequency calls |
| 中文支持 | 7.5 / 10 | Benchmarked across multiple languages, but official Chinese-specific data is not disclosed separately |
| 输出质量 | 9.0 / 10 | Highest accuracy across 36 blind-test benchmarks—strong robustness |
Overall rating: 8.7/10
Frequently Asked Questions (FAQ)
What is Microsoft-Decision-1?
It is a decision model released by Microsoft in October 2026, specifically designed for structured decision tasks—including routing, classification, priority ranking, and result validation—rather than text generation. Built via post-training on Qwen3.5-9B, it emphasizes speed, low cost, and high accuracy.
Is Microsoft-Decision-1 free? How is it priced?
Input tokens cost $0.042 per million tokens; output tokens are completely free. Compared to general-purpose LLMs that charge for both input and output, this reduces costs by one to two orders of magnitude for high-frequency decision tasks.
How do I integrate Microsoft-Decision-1?
Two options: first, call it directly from the model catalog on Microsoft Foundry (ai.azure.com); second, access it via the OpenRouter platform. Both require only a single structured API request—pass in your options and receive probability scores.
Decision models and GPT-6Claude What distinguishes these large models?
Large models handle generation and complex reasoning; decision models only perform “making judgments among given options.” Decision models are dozens of times faster and hundreds of times cheaper, but cannot write articles or code. They operate in a division-of-labor relationship—large models do the work, while decision models oversee, validate, and orchestrate—commonly collaborating within Agent systems.
What use cases is it suited for?
Best suited for high-frequency, low-latency scenarios requiring structured outputs: content moderation and classification, customer support ticket routing, AI response quality assurance, user feedback tagging, and next-step decisions within Agent workflows. Not suitable for open-ended content generation or long-chain reasoning.
Want to learn more about new models and tools? Browse our AI Model Library, or use Tool Comparison Engine model selection by use case. Related reading:GPT-6.1 Sol Deep Review · GPT-6 Sol vs. Luna Evaluation · Strands Decider 2B Decision Model.

Pingback: OpenAI Releases Nearly 400 Mathematical Results at Once: 700 Manuscripts, Mathematicians Call It Pure Madness – AI Dash