It doesn't answer. It scores.

Microsoft's new model has no interest in conversation. Decision-1 takes a fixed set of options and returns a calibrated probability for each: yes or no, multiple choice, rating scales, or rubric-based grades of AI responses and agent actions. The probability is the product. A 90% score is supposed to mean the model is right nine times out of ten, and downstream software uses that confidence to decide whether to act, defer, retry, escalate, or hand the case to a human.

Microsoft is selling this as the control layer for agentic systems: routing, classification, prioritization, verification, data labeling, AI judging, safety screening. The Xbox Research team used it to sort more than 10,000 game feedback items into researcher-defined themes. The Copilot team uses it to grade chat responses. Microsoft Discovery uses it for adaptive replanning in experiments.

The Qwen part is the story

Decision-1 is post-trained from Qwen3.5-9B, Alibaba's open-weight 9-billion-parameter model. Microsoft says it plans to rebase on other models later, including its own technology and OpenAI's. For now, Microsoft's newest model runs on a Chinese lab's weights.

Read that against the pricing. Decision-1 costs $0.042 per million input tokens with output free. OpenAI's Luna Decisions, the closest competitor, charges $0.10 per million with free output. Microsoft, OpenAI's largest backer, is undercutting OpenAI's own decision model by 58 percent using Alibaba's base. The category itself is filling in fast. We covered Cloudflare's Clef decision models nine days ago; H2O.ai ships H2O-Lightning-4B. The pattern is a model that doesn't talk. It decides.

Where the numbers came from

Every performance figure in the announcement comes from Microsoft's own chart, so read accordingly. The 83.5% accuracy is an average across 36 benchmarks covering nearly 150,000 questions that Microsoft says were withheld from training. Quyet-1.0-Large scored 81.9% on the same chart, H2O-Lightning-4B 77.2%, OpenAI's Luna Decisions 79.4%. None of this has been independently reproduced.

The latency table is worse. Microsoft reports its own model at 85ms p50, measured through its own Foundry service, then lists competitor figures taken from JevBench's "adjusted median" — a methodology JevBench itself says doubles measured time and adds 0.15 seconds. H2O.ai disputes its entry directly: its measured median is 29ms, not the 210ms on Microsoft's chart. And Microsoft-Decision-1 does not appear on the JevBench board at all, where H2O-Lightning-4B holds the top composite score.

The internal trials — Xbox sorting 10,000 feedback items "14 times faster and 200 times cheaper than GPT-6 Sol," Copilot grading "100 times faster," Discovery's "46 times more consistent" judging — are Microsoft measuring itself and reporting the results. Useful as a directional signal, useless as a scoreboard.

Who should actually care

Teams running agents at volume. If you route requests between models, grade outputs, or need a cheap judge for evals, a calibrated 85-millisecond scorer at $0.042 per million input tokens is genuinely useful infrastructure. The price is the real number in this announcement. The benchmark table is a brochure.

Sources

  1. [1] TestingCatalog — “Microsoft launches Decision-1 model in Foundry”Read source
  2. [2] MarkTechPost — “Microsoft AI Releases Microsoft-Decision-1”Read source
  3. [3] AI Weekly — “Microsoft ships Decision-1”Read source
  4. [4] Analytics Insight — “Microsoft Decision-1: Faster, Cheaper AI Model”Read source
  5. [5] Windows Forum — “Microsoft Decision-1 in Foundry”Read source
  6. [6] AI Tool Herald — “Microsoft releases Decision-1 routing and classification model”Read source