On October 9, 2026, Microsoft introduced Microsoft-Decision-1, a new model for fast decision scoring. Unlike traditional large language models (LLMs), it does not generate text: the model takes a strictly defined set of options and returns a calibrated probability score for each option through a structured API. The new model is available on the Microsoft Foundry platform and via the OpenRouter service.

How the Model Works

Microsoft-Decision-1 was built specifically for the decision-scoring task. Instead of a free-text prompt, it receives a ready-made list of options from the user and assigns each a probability score — for example, yes/no answers, multiple-choice options, or ratings on a scale. In addition, the model can score AI responses and AI agent actions against predefined criteria (rubrics).

"Unlike text-generating LLMs, Microsoft-Decision-1 accepts a strictly defined set of options and returns a calibrated probability score for each option via a structured API; it supports yes/no, multiple-choice and rating options, as well as criterion-based scoring of AI responses and agent actions." — Microsoft official announcement

Microsoft is targeting the model at tasks designed specifically for agentic systems: routing, classification, prioritization, verification, and workflow management. These are all small but critical decisions that modern AI agents make every second.

What It Is Built On

Microsoft created the model through additional training (post-training) on top of Alibaba's open-weight Qwen3.5-9B model. Training in a single pass was focused on the decision-scoring task. The company said it plans to move the model to Microsoft AI (MAI) and OpenAI models as its base in the future.

This approach shows a path of specializing an existing open model instead of training a large language model from scratch. Such post-training lets a smaller model match large models on a narrow task. The choice of Qwen3.5-9B is no accident: specializing on an open-weight base gives Microsoft full control to adapt the model to its own infrastructure and the Foundry ecosystem. And the plan to move the base model to MAI and OpenAI models in the future signals that the company sees this direction as a long-term strategy.

Pricing and Availability

The model is available on the Microsoft Foundry platform and via the OpenRouter service. These two channels let developers quickly plug the model into existing workflows: Foundry suits those working inside Microsoft's cloud ecosystem, while OpenRouter suits those using models from multiple providers through a single API. The pricing — $0.042 per million input tokens, output tokens free — shows the model is designed precisely for tasks that process many small requests: no large text is generated per request, only a score is returned.

Speed and Accuracy

In its tests, Microsoft evaluated the model across 36 benchmarks based on nearly 150,000 questions not shown during training. According to the company, Microsoft-Decision-1 recorded the highest accuracy of all models in the test.

The speed results are also notable: the model ran 2.5 times faster than the runner-up H2O-Lightning-4B v1.1 and 35 times faster than GPT-6 Sol on the P50 latency metric. P50 is a median measure meaning more than half of requests get a response faster than this time — so the gap is not random but stable. The pricing matches: $0.042 per million input tokens, with output tokens free. In other words, the model is positioned as an inexpensive service for tasks that mostly score incoming requests. The combination of high accuracy and low latency matters precisely for real-time agent systems: an agent sends dozens of such scoring requests at each step, and every millisecond adds to the total response time.

Safety and Reliability Testing

Microsoft safety-tested the model on 5,250 requests across 11 benchmarks covering harmful content, jailbreak attempts, and prompt injection attacks. In stability tests, small changes to input data changed decisions in only 1.3% of cases on average. This indicates the model's resilience to noisy or deliberately corrupted inputs.

Xbox Research in Practice

As an example of practical use, Microsoft cites the experience of its Xbox Research team. The team tested Microsoft-Decision-1 to analyze more than 10,000 user reviews of games. The model's quality turned out to be competitive with GPT-6 Sol, while running more than 14 times faster and 200 times cheaper. This shows the advantage of a specialized small model in high-volume repetitive scoring tasks. Analyzing game reviews is exactly such a task: thousands of short texts need quick classification or priority sorting, where generating a full response for each text would be wasteful. Microsoft's use of this example signals that the target audience is teams building large-scale automated evaluation pipelines.