AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Picking an AI model by its brand name may be a costly shortcut. In Firmulate’s latest company-running trial, Moonshot’s Kimi K3 placed second, ahead of three of four Western frontier models. The result points to a practical question for businesses adopting AI agents: how does a model handle pressure, follow through, and act on what it finds?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A tougher test than a chat demo

Firmulate put each frontier model in charge of the same small software company during its worst week. The models faced the same customers, crises and temptations, and every decision was versioned and auditable. The experiment is live and watchable at Firmulate.

The final July 2026 league table puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Finding the detail—and closing the deal

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap between understanding the job and completing it is central to the trial: “Same diagnosis, same pitch — no signature.”

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. Kimi K3 found that buried fact, won the deal, saved the churning customer and resisted all three baits. It made one deviation, the fewest in the field.

The baits included fake CEO messages that escalated over three stages and a reporter asking for “just one yes/no, on background.” All five models refused. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as execution

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. Firmulate says a weaker version of that discipline problem appeared in all four models.

The company behind the test has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules and every workday versioned. Those details make the trial concrete, but the scores describe performance in this particular simulated company, not a guarantee about every business.

Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice. Enterprises can run the same wargame against a read-only export of their own business; nothing writes back to real systems. Results and details are available on the Firmulate benchmarks page.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Test the job, not just the model

Kimi K3’s second-place finish shows that the league is open: it beat Sonnet 5, Fable 5 and Opus 4.8, while trailing gpt-5.6-sol by two points. The useful lesson for companies is to test models against the work they will actually do. A polished answer is not the same as finding a buried fact, respecting boundaries and finishing the sale.

Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ron Gilbert Started Production On Thimbleweed Park 2

Game designer Ron Gilbert has started production on Thimbleweed Park 2, confirming a sequel to his acclaimed adventure game. Details are still emerging.

Train Sim Created By Just One Person Is Being Called The Best Ever Made

A solo developer’s train simulation game is being hailed as the best ever made, garnering widespread acclaim for its quality and detail.

Copyright Challenges for Remix Culture in 2025

Struggling with legal uncertainties and technological barriers, remix culture in 2025 faces complex copyright challenges that demand careful navigation to succeed.

Electronic Drum Kits Need Thoughtful Placement to Stay Apartment-Friendly

Meta description: “Mastering your electronic drum kit placement helps keep noise down in apartments—discover essential tips to create a quieter practice space.