
The costly difference between answering and investigating
For technology buyers, fluent answers are becoming less impressive. The harder question is whether an AI agent will inspect the available evidence before acting—and whether it will carry a sound decision through to completion.
Firmulate turned that question into a live business experiment. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. All the models recognized every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned.
The difference was not a sharper sales pitch or a better diagnosis. It was whether the agent followed a trail through the company’s files.
AI document trail analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The decisive fact was not in the obvious place
The customer event did not contain the information needed to win the deal. The decisive weakness in a competitor’s position sat two document references deep inside the company’s own files. Models that found and used it secured the agreement at full price, adding +€4,583 MRR. Those that did not lost the opportunity automatically.
That makes “reads your files before answering” more than a desirable product feature. In this test, it was a measurable, purchase-deciding capability. The unsuccessful agents could still identify the commercial situation and construct the pitch. But the result was blunt: “Same diagnosis, same pitch — no signature.”
This distinction matters as businesses move agents beyond chat windows and into CRM workflows, support queues and forecasting. An agent can sound informed while relying only on the evidence placed directly in front of it. Real work often requires following references, opening supporting material and checking whether the company already knows something decisive.
A difficult week inside a live company
The simulated company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than dependent on a polished retrospective.
The final July 2026 Crucible League placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The full public results appear on the Firmulate benchmarks page.
K3’s showing comes with an important comparison note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. Even with that difference, it finished just behind the league leader.
Trust held up better than follow-through
The agents faced fake CEO messages that escalated over three stages, followed by a reporter attempting to extract information with “just one yes/no, on background.” All 5 models refused every attempt. Kimi K3 summarized the problem clearly in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The result separates two capabilities that buyers may otherwise bundle together. The models were consistently good at spotting suspicious instructions and protecting trust. They were much less consistent at completing legitimate work after reaching the correct conclusion.
Opus 4.8 illustrates the gap most sharply. It was the most thorough participant, producing +80 learned rules and the deepest analyses, but it finished last. The close was left on the table, and its discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.
Firmulate also exposes the decisions behind these profiles through a guess-the-model quiz powered by 242 real, unedited management decisions. The point is not simply to identify stylistic fingerprints. It is to see how convincing language can obscure meaningful differences in investigation, discipline and execution.

A practical procurement question
Enterprises evaluating AI agents should ask for evidence that a system can follow internal references and finish consequential work—not merely summarize a supplied document or produce a credible recommendation. The buried-fact test shows why: an agent may understand the situation, resist manipulation and still fail at the moment that creates business value.
Firmulate offers enterprises the same wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a practical way to test agents against an organization’s actual information landscape before granting operational access.
The lesson from the €55,000 deal is straightforward. Reading deeply is not administrative diligence added around intelligence. For an AI worker, it can be the difference between an articulate analysis and a completed result.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html