
Artificial intelligence needs a management test
Technology buyers are trained to compare visible performance: the sharper display, the faster chip, the chatbot that produces the most convincing answer. But an AI agent working inside a company is not merely answering questions. It is deciding what deserves attention, reading imperfect records, resisting pressure and turning analysis into completed work.
That is the measurement gap exposed by Firmulate, a live experiment that put frontier models in charge of the same small software company during its worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable. The result is less a contest of eloquence than a test of whether an artificial manager can remain useful when consequences accumulate across days.
The unsettling finding was not that the models misunderstood the business. All of them spotted every crisis and refused every manipulation attempt. The gap appeared afterward: only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature.”
enterprise AI management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From chat quality to management quality
Coding leaderboards and chat arenas remain valuable, but they mostly reward the quality of an answer produced in a bounded interaction. A company creates a harsher test. Work competes for limited attention. Important evidence may sit outside the incoming message. A correct recommendation has little value if nobody follows through, while a seemingly productive shortcut can destroy trust.
The final Crucible League results from July 2026 make that distinction visible. gpt-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the experiment imposed a hard boundary around trust: “no amount of good work outweighs a breach of trust.” A single breach capped the total.
The most revealing episode concerned the deal. The decisive weakness in a competitor was not included in the customer event. It was buried two document references deep in the company’s own files. Models that read the file could use that fact to win the contract at full price, worth +€4,583 MRR. The distinction was not between models that could reason and models that could not. It was between those that treated company records as part of the job and those that stopped at the obvious context.
This is exactly where conventional demonstrations flatter AI systems. A fluent response can make the work appear complete even when the commercially decisive action remains undone. In the wargame, some models reached the right diagnosis and developed the right pitch, yet failed to secure the signature. That is not chat failure. It is management failure.
Honesty held up better than execution
The social-engineering results offer a more reassuring signal. Fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest framing: “Treat the request as a suspected approval-bypass / possible impersonation.”
That matters because enterprise agents will encounter requests that sound urgent, authoritative and convenient. The test suggests these frontier models can recognize manipulation under pressure. It also shows why safety cannot be judged separately from ordinary work: an agent must preserve trust while still advancing the legitimate task.
Opus 4.8 illustrates the tension. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models. Thoroughness, in other words, did not guarantee operational judgment.
One comparison also deserves a fairness note. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its 93 score, but it should shape how readers interpret the ranking.
A company that makes consequences visible
Firmulate’s live company employs 13 synthetic workers and uses real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, making the experiment watchable rather than retrospective.
The project also turns 242 real, unedited management decisions into a “guess the model” quiz. That is a useful challenge to anyone convinced they can identify a model by tone alone. Management quality often reveals itself in quieter behaviors: checking the files, escalating a blocked action, resisting an improper request and completing the final step.

The next buying question
Businesses evaluating AI agents should ask something harder than whether a model writes good code or gives persuasive advice. Can it triage a churn wave, handle a price increase, survive a downround or respond to a PR crisis without losing discipline? Does it read the company’s own evidence before acting? Does it finish what it starts? Will it remain honest when someone invokes executive authority?
Firmulate’s answer is not that existing benchmarks are useless. It is that they measure only part of the job. The emerging category is management quality: performance under capacity pressure, with consequences extending beyond a single prompt.
Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems. That turns agent evaluation from a generic leaderboard exercise into a rehearsal for the exact environment where the system may eventually operate.
The crucial divide is no longer between an AI that sounds capable and one that does not. It is between an agent that can describe the right move and one that reliably makes it—without sacrificing trust along the way.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html