
Picking an AI model by its brand name may be a costly shortcut. In Firmulate’s latest company-running trial, Moonshot’s Kimi K3 placed second, ahead of three of four Western frontier models. The result points to a practical question for businesses adopting AI agents: how does a model handle pressure, follow through, and act on what it finds?
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A tougher test than a chat demo
Firmulate put each frontier model in charge of the same small software company during its worst week. The models faced the same customers, crises and temptations, and every decision was versioned and auditable. The experiment is live and watchable at Firmulate.
The final July 2026 league table puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
As an affiliate, we earn on qualifying purchases.
Finding the detail—and closing the deal
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap between understanding the job and completing it is central to the trial: “Same diagnosis, same pitch — no signature.”
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. Kimi K3 found that buried fact, won the deal, saved the churning customer and resisted all three baits. It made one deviation, the fewest in the field.
The baits included fake CEO messages that escalated over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.
The company behind the test has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules and every workday versioned. Those details make the trial concrete, but the scores describe performance in this particular simulated company, not a guarantee about every business.
Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Results and details are available on the Firmulate benchmarks page.

Test the job, not just the model
Kimi K3’s second-place finish shows that the league is open: it beat Sonnet 5, Fable 5 and Opus 4.8, while trailing gpt-5.6-sol by two points. The useful lesson for companies is to test models against the work they will actually do. A polished answer is not the same as finding a buried fact, respecting boundaries and finishing the sale.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
