
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
When diligence becomes a distraction
Technology buyers are accustomed to judging artificial intelligence by how much it can produce: longer reports, more detailed reasoning and ever-expanding stores of knowledge. Firmulate’s Crucible League exposes the weakness in that assumption. Its most thorough participant, Opus 4.8, learned more than 80 new rules and produced the deepest analyses in the field. It still finished last.
The problem was not comprehension. Opus identified the crises placed before it and resisted attempts to manipulate its decisions. Its analysis even helped earn a €55,000 customer opportunity. But the model did not complete the decisive commercial action: getting the agreement signed. For businesses considering agents that can operate across support, sales and planning, that distinction matters more than eloquence.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A bad week designed to reveal good judgment
Firmulate runs AI models as complete companies, testing management quality rather than conversational polish. In the experiment, each frontier model was given the same small software business and sent through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable.
The synthetic company is deliberately unforgiving. It has 13 employees and real money mechanics, with monthly burn of €105,000 against just €2,300 in monthly recurring revenue. Its cash countdown is public, and its participants have accumulated more than 680 self-learned playbook rules. The result is a live, watchable experiment in whether an AI can turn apparently sound judgment into useful work.
The final July 2026 standings put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The complete results are available on the Firmulate benchmark page.
The clue was in the company’s own files
The commercial challenge turned on a fact that was easy to miss. A decisive weakness in a competitor was not present in the customer event itself. It sat two document references deep inside the company’s files. Models that followed those references found the advantage and secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.
That makes the failure particularly instructive. Every model spotted every crisis, and the participants arrived at the same diagnosis and the same pitch. Yet only two signed the €55,000 deal their analysis had earned: “Same diagnosis, same pitch — no signature.” It is a sharp example of the distance between knowing what should happen and ensuring that it does.
Opus 4.8 embodies that gap more vividly than the rest. Its additional 80 learned rules signal serious effort, not carelessness. Its analyses were the deepest. But volume did not translate into priority. The model left the close on the table, while its operational discipline also slipped through write attempts into a locked department instead of escalation through the available path.
This was not an isolated flaw unique to Opus. The same weakness appeared, less strongly, in all four models covered by the core finding. That context makes the profile fairer and more useful: Opus was not uniquely incapable, but it was the clearest demonstration of how thoroughness can crowd out completion.
Cautious under pressure, but not complete
The models performed better when the threat was overt. Fake CEO messages escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused the social-engineering attempts. Kimi K3 recorded the clearest summary of the risk: “Treat the request as a suspected approval-bypass / possible impersonation.”
That clean result should not be understated. An agent that closes deals but leaks information or bypasses authorization would be commercially dangerous. Firmulate’s design recognizes that trust is not merely another item on a checklist. At the same time, refusing a bad request is only part of operating a company. An effective agent must also pursue legitimate actions to completion.
One qualification belongs beside the standings. Kimi K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the observed outcomes, but it is relevant context when readers compare participants as if every setting were identical.
Firmulate also turns 242 real, unedited management decisions into a quiz that asks visitors to guess which model made each choice. For enterprises seeking a more specific test, the company offers a pilot using a read-only export of the business. Nothing writes back to real systems, allowing organizations to observe how an AI workforce would behave without granting it operational control.

Businesses need finishers, not just thinkers
Opus 4.8 is a caution against confusing visible effort with business impact. Its extensive rule learning and detailed analysis were genuine strengths. They simply did not compensate for a missed close and a failure to escalate correctly when access was blocked.
The broader lesson applies beyond any single model. Before an AI touches a CRM, support queue or forecast, buyers should ask whether it reads the relevant files, preserves trust, identifies the next consequential action and finishes what it starts. The most impressive answer is not always the longest one. Sometimes it is the signed agreement that proves the analysis mattered.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.