AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

AI’s next big test is not conversation. It is judgment.

Technology buyers have become accustomed to comparing artificial intelligence through polished answers, coding demonstrations and benchmark charts. Firmulate poses a more consequential question: when several frontier models face the same troubled company, can you tell which one will actually behave like a capable manager?

Its interactive guess-the-model quiz draws on 242 real, unedited management decisions. Readers encounter the models through what they chose to do—not through marketing claims—and then try to identify the decision-maker. The result is part quiz, part management case study and part warning about evaluating AI through chat alone.

The decisions come from a live experiment in which each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable. That controlled setup exposed recognizable differences in thoroughness, discipline and follow-through: what might reasonably be called management personalities.

Amazon

AI decision-making management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Identical crises, sharply different outcomes

The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One boundary was absolute, however: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The striking finding was not that some models missed obvious emergencies. They did not. Every model spotted every crisis and refused every manipulation attempt. The separation appeared later, at the point where analysis had to become action. Only two signed the €55,000 deal their own work had earned. As Firmulate summarizes the gap: “Same diagnosis, same pitch — no signature.”

That is an unusually useful distinction for businesses considering AI agents. Recognizing a problem, drafting a convincing response and completing the commercially important step are different capabilities. A fluent answer can disguise the distance between understanding the work and finishing it.

The clue hidden in the company’s own records

The decisive weakness of a competitor was not sitting in the customer event. It was buried two document references deep in the company’s files. The models that followed that trail won the deal at full price, worth +€4,583 MRR.

For quiz players, such episodes make authorship more interesting than a hunt for writing tics. A long answer may reveal a model’s appetite for investigation, but detail alone does not guarantee the best outcome. A terse decision can be disciplined or incomplete. The revealing evidence lies in which facts the model seeks, which instructions it resists and whether it carries a promising opportunity through to completion.

Pressure did not break the trust boundary

The social-engineering tests escalated through fake CEO messages over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly crisp assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

This shared resistance matters because the simulated company was not an abstract puzzle. Firmulate describes a business with 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and every workday versioned. The live experiment is watchable, allowing the public to examine the operating behavior behind the results.

When thoroughness becomes a trap

Opus 4.8 produced the deepest analyses and added +80 learned rules, making it the most thorough participant. Yet it finished last. It left the close on the table, while discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in all four other participants, though less strongly.

That profile complicates the popular assumption that more reasoning automatically means better management. Opus 4.8 generated abundant evidence of effort, but the league rewarded useful completion and disciplined conduct. Kimi K3, meanwhile, requires a fairness note: it ran with the API default and without an effort parameter, while the others ran at xhigh. That difference should remain visible when readers interpret the standings.

Infographic —
The findings at a glance — source: firmulate.com.

A personality test with operational consequences

The quiz works because the choices are entertaining to compare, but its deeper lesson is practical. Frontier models can agree about the nature of a crisis and still diverge at the moment of execution. They can share strong defenses against manipulation while differing in research habits, procedural discipline and commercial follow-through.

For gadget enthusiasts and technology leaders alike, that changes what “best model” should mean. The decisive question is not merely which system sounds smartest. It is which one reads the relevant files, protects trust under pressure and completes the valuable work already within reach.

Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business, with nothing written back to real systems. That makes the public experiment more than a spectator sport: it demonstrates why an AI workforce should be tested in realistic operating conditions before it receives consequential responsibilities.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Gta Trailer

Rockstar Games has officially released the first trailer for Grand Theft Auto 6, confirming key details and setting the stage for the upcoming game launch.

In‑Theater Immersive Sound Systems Explained

AIThis post was created with the assistance of artificial intelligence (AI).In-theater immersive…

Netflix’s Horror Game ‘Unhinged’ Is Funny, Scary and Points to a Future Without Second Screens

Netflix releases ‘Unhinged,’ a horror game praised for its mix of humor and scares, signaling a potential shift in interactive entertainment without second screens.

The Return of Midnight Movie Premieres: Nostalgia or Necessity?

Cultural nostalgia meets cinematic necessity as midnight movie premieres revive communal experiences—discover how this trend balances tradition with modern demand.