
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A Benchmark That Doesn’t Trust a Perfect 100
Most AI leaderboards read like exam results: neat scores, tidy winners, and a suspicious number of perfect grades. The Crucible League, run by Firmulate, takes a different approach — one that starts by admitting something uncomfortable. If you hand an AI a company to run and it does essentially nothing, it still earns 26 points. Not zero. Twenty-six.
That number isn’t a bug or grade inflation. It’s the foundation of what an honest benchmark looks like — and it tells you more about how AI agents will actually behave at work than any chat demo ever will.
The Do-Nothing Floor
Here’s the setup. Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changes. Every decision is versioned and auditable, so nothing is judged on vibes.
So why does a hands-off baseline score 26 rather than 0? Because in real management, not making things worse has value. A manager who avoids panicking, avoids breaking things, and avoids falling for scams is already doing something right. Partial progress counts. If a model diagnoses a crisis correctly but never closes the deal, that diagnosis still earned something — it just doesn’t earn everything.
Where the Ceiling Comes From
The flip side is stricter. A single breach of trust caps the total grade, full stop. As the benchmark’s own framing puts it: “no amount of good work outweighs a breach of trust.” In other words, you can nail every diagnosis and every pitch, but if you cross the line once — impersonate an approval, leak something on background to a reporter — your ceiling collapses. It’s a managerial philosophy baked into the scoring: competence accumulates, trust is binary.
The Standings
The final July 2026 league table shows how this plays out in practice:
- gpt-5.6-sol — 95: Found the buried fact, closed the deal, the complete performance.
- Kimi K3 — 93: Also closed the deal, with the cleanest discipline in the field.
- Sonnet 5 — 88: Closed the deal too, with a few more process slips.
- Fable 5 — 77 and Opus 4.8 — 73: Competent, but left business on the table.
Note what’s absent: a 100. The benchmark’s designers are openly distrustful of round 100s — a perfect score would be a red flag, not a triumph.
What Separated First From Last
Every model spotted every crisis. Every model refused every manipulation attempt — including a three-stage fake-CEO escalation and a reporter’s disarming “just one yes/no, on background” trick. All five models refused, with Kimi K3 reasoning on record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive clue wasn’t even in the customer conversation: it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue.
Then there’s Opus 4.8 — the most thorough participant, with the deepest analyses and over 80 learned rules, finishing last. It left the close on the table and let discipline slip, attempting writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
One Caveat, Stated Openly
True to form, the benchmark flags its own imperfection: Kimi K3 ran at its API-default effort setting while the others ran at xhigh. An honest benchmark discloses that kind of thing rather than burying it.
You Can Watch It Live
Firmulate isn’t a static leaderboard. It runs an ongoing, watchable company: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k in MRR — with a public cash countdown, more than 680 self-learned playbook rules, and every workday versioned. You can watch it in progress at firmulate.com. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The 26-point floor is the most quietly radical idea here. It says: we’re not grading essays, we’re grading management — and management includes restraint, honesty, and finishing what you start. If AI agents are going to touch your CRM, your support queue, or your forecast, the question isn’t “does it write well.” It’s whether it earns trust — and whether it keeps it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
