The motivating example
The paper opens with a support agent handling a late delivery. It makes nine tool calls, reads the refund policy correctly, opens a support ticket, and closes it as resolved. A grader checking tool calls or the final sentence would call this a success. The database disagrees: the courier exception is still open, so the ticket should have been left on hold, and the customer never got an answer to her actual question. One field, the ticket status, has the wrong value. The agent's fluent explanation cannot repair that state.
The uncomfortable arithmetic
The authors ran a common-set ablation: 121,680 valid trials across 12 models. Of the 79,853 attempts that failed the checks, 67.24 percent terminated cleanly, with no reported tool error. They looked fine. They were not. Among the failures, 77.61 percent had wrong field values, 43.30 percent produced unintended extra effects, and 25.36 percent were missing a required effect. Most of the breakage is in bookkeeping, not reasoning.
Then the distinction that matters for anyone shipping agents. Claude Opus 5.5 leads on single-attempt accuracy at 67.16 percent. But run each task 20 times and the story changes: Opus 5.5 passes the exact same number of tasks on all 20 attempts as Opus 5, 241 of 507. The accuracy gain bought no added dependability. Kimi-K3 is the starkest case: the strongest open-weight model solves 476 of 507 tasks at least once, 93.89 percent, yet passes every attempt on only 68 tasks, 13.41 percent. Solving a task once is a capability. Solving it every time is a product.
The price of dependability
The cost arithmetic is worth a second read. Per successful attempt, GPT-5.6 Sol is the cheapest of the models at an estimated $0.127. But most deployments are not paying per success, they are paying for repeatability. On cost per task that passed all 20 attempts, Claude Opus 5.5 comes in at $7.80 against $9.76 for GPT-5.6 Sol. The cheapest model per lucky run is not the cheapest model per reliable run. Anyone budgeting an agent deployment on per-token pricing is doing the wrong math.
Read the release note
Some honesty is in order, and to its credit the release invites it. These figures describe the authors' test environment, not your production database. The 507 tasks are synthetic reconstructions, the policies are fixed, and nothing here shows the benchmark predicts performance against live company records. The research paper has been on arXiv since August 20, so October 3 is a release date for a runnable benchmark, not a discovery date for new science. And the authors say they have not measured whether the obvious fixes, targeted retries, human approval on irreversible changes, improve scores. The benchmark tells you where the failure is. It does not tell you how to fix it.
What it does tell you is that the industry's eval stack has been grading the wrong thing. A transcript that reads well and a tool trace without errors are exactly the two signals demos are built around. ThinkingBox makes the case that neither is the work. The work is the record left behind. For anyone deciding whether an agent is ready to touch customer accounts, bookings, or claims, that is the right question, and until now there was no standard way to ask it.
Sources
- [1] RuntimeWire — “Microsoft releases an agent benchmark where the database gets the final vote” (Oct 3, 2026)Read source
- [2] FourWeekMBA — “ThinkingBox: Opus 5.5 and Opus 5 Tie at 241 of 507 Tasks” (Oct 3, 2026)Read source
- [3] DEV / mikefluff — “New Benchmark Catches AI Agents Lying About Finished Work”Read source
- [4] Hugging Face dataset — microsoft/ThinkingBox-BenchRead source
- [5] GitHub — microsoft/thinkingbox-dataRead source