AI coding tools
Microsoft-Decision-1: A $0.042 Decision Model for Agent Loops
Short answer
- Microsoft launched Microsoft-Decision-1 on : a "decision model" that reads a situation plus a fixed set of options and returns a calibrated probability for each option — JSON out, no prose. It's post-trained from Alibaba's Qwen3.5-9B, with a 32,768-token context window and hosted-API-only delivery (no open weights).
- Pricing: $0.042 per million input tokens, output tokens free — the exact structure TypeSafe's Jev has used since mid-September.
- Microsoft's own 36-benchmark comparison (147,137 questions): 83.5% average accuracy, 85 ms p50 latency — 35× faster than GPT-6 Sol (3.01 s) on the same harness. Every number is vendor-run; nothing independent yet.
- Best fit: routing, classification, verification, and agent-loop control — thousands of cheap calls per task, not conversation.
Most model launches ask you to read more text. This one asks you to stop generating it. Microsoft-Decision-1, announced on Microsoft's Command Line blog and in the Azure AI Foundry catalog, is a small model built to do one thing: take a situation, a question, and a fixed list of answer options, then return a probability for each option in a single pass. Yes/no gates, multiple-choice routing, rubric grading of AI outputs, groundedness checks against supplied evidence — with an explicit abstention option ("cannot tell") so the model can decline rather than guess. Output is JSON, and it never writes explanations.
Microsoft is framing this as a new category — the decision model — and the framing matters more than the marketing. Agent pipelines are full of moments where you don't need a paragraph; you need a confident pick from five options, made cheaply, thousands of times per run. That's the job Decision-1 is shaped for.
The pricing math, and who it copies
$0.042 per million input tokens. Output unmetered. If that structure feels familiar, it should: TypeSafe's Jev launched with the same numbers in September — $0.042/M in, output free, a rate datacamp and others described as "too cheap to meter." Microsoft's decision model lands 23 days later with an identical price sheet. Whether that's convergence or imitation, the effect for buyers is the same: the expensive part of high-frequency scoring calls — output tokens — now costs nothing on both.
What does that buy you in practice? A router making 1,000 calls a day with ~500-token prompts burns roughly 15 million input tokens a month — about $0.63 at Decision-1's rate. The same loop on a general-purpose model with typical input/output pricing runs orders of magnitude higher, mostly because generated routing text isn't free. Illustrative math, obviously — your prompt sizes will differ — but the shape of the curve is the point: per-call costs collapse to rounding errors, and architecture decisions stop being pricing decisions.
The benchmarks are Microsoft's own — read them accordingly
The launch numbers come entirely from Microsoft's harness: 9 systems, 36 benchmarks, 147,137 questions kept blind from training. On that harness Decision-1 averaged 83.5% accuracy, ahead of Quyet-1.0-Large at 81.9%, though it placed second on calibration (92.2 vs 93.1). Latency is the sharper claim: 85 ms p50 and 125 ms p95 — 4.5× quicker than Quyet-1.0-Large and 35× quicker than GPT-6 Sol, which needed 3.01 seconds on the same test. Microsoft also perturbed each request 8 ways (paraphrasing, option shuffling, and kin) and reports the model flipped its decision on just 1.3% of them.
None of this is independently verified. MarkTechPost's own launch write-up flags the same thing: the best numbers are real on Microsoft's harness, and the worst fact is that every benchmark is vendor-run. Treat 83.5% as "Microsoft says 83.5%" until someone else runs the model against their own workload — which, given the $0.042 entry price, costs approximately nothing to do.
What's under the hood — and what's not public
The base is Qwen3.5-9B, post-trained by Microsoft for single-pass decision scoring — trained on public datasets under Microsoft's Open Data process plus synthetic data. No open weights, no quantized variants, no hardware details. Microsoft says future versions will rebase on MAI and OpenAI models, and — per OpenRouter's notes quoted in launch coverage — weights update continually while the API shape stays fixed.
One availability discrepancy worth knowing: launch coverage describes Decision-1 as available in Microsoft Foundry and OpenRouter. Foundry checks out — the model has a card in Azure's catalog. But when we queried OpenRouter's public model index on — twice — its 458 listed models didn't include Decision-1. Maybe it's rolling out in stages, maybe the listing lags. If OpenRouter access is your plan, verify before you build on it.
Should you route your agents through it?
If your pipeline burns general-model tokens on classification, verification, or tool-selection steps, the math here is hard to ignore — Decision-1 answered 35× faster than GPT-6 Sol on Microsoft's own harness, and output-free pricing removes the biggest variable in loop-heavy workloads. We compared how AI coding tools structure their own usage pricing in our Cursor vs CodeBuddy breakdown, and the pattern holds across the market: the per-decision cost, not the per-token sticker price, is what your bill actually tracks.
The honest reservations: a 9B model has a ceiling on nuanced judgment, the abstention behavior ("cannot tell") needs testing on your own edge cases before you trust it, and you're buying into a vendor-defined category where the only published numbers are the vendor's. The counterargument is built into the price — at $0.042 per million input tokens, running Decision-1 against your own labeled data for a day costs less than the coffee you'll drink doing it. Verify on your workload, then decide.
Sources: MarkTechPost — Microsoft AI Releases Microsoft-Decision-1 (Oct 9, 2026) · Microsoft Command Line blog launch post ($0.042/M input, free output; named — bot-walled, content corroborated via search index) · Azure AI Foundry model catalog (32K single-pass scoring description; named). Jev pricing comparison per TypeSafe's model documentation as reported by DataCamp and MindStudio, September 2026. Checked .