AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

When diligence becomes a distraction

Technology buyers are accustomed to judging artificial intelligence by how much it can produce: longer reports, more detailed reasoning and ever-expanding stores of knowledge. Firmulate’s Crucible League exposes the weakness in that assumption. Its most thorough participant, Opus 4.8, learned more than 80 new rules and produced the deepest analyses in the field. It still finished last.

The problem was not comprehension. Opus identified the crises placed before it and resisted attempts to manipulate its decisions. Its analysis even helped earn a €55,000 customer opportunity. But the model did not complete the decisive commercial action: getting the agreement signed. For businesses considering agents that can operate across support, sales and planning, that distinction matters more than eloquence.

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week designed to reveal good judgment

Firmulate runs AI models as complete companies, testing management quality rather than conversational polish. In the experiment, each frontier model was given the same small software business and sent through its worst week. The customers, crises and temptations remained constant, while every decision was versioned and auditable.

The synthetic company is deliberately unforgiving. It has 13 employees and real money mechanics, with monthly burn of €105,000 against just €2,300 in monthly recurring revenue. Its cash countdown is public, and its participants have accumulated more than 680 self-learned playbook rules. The result is a live, watchable experiment in whether an AI can turn apparently sound judgment into useful work.

The final July 2026 standings put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” The complete results are available on the Firmulate benchmark page.

The clue was in the company’s own files

The commercial challenge turned on a fact that was easy to miss. A decisive weakness in a competitor was not present in the customer event itself. It sat two document references deep inside the company’s files. Models that followed those references found the advantage and secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.

That makes the failure particularly instructive. Every model spotted every crisis, and the participants arrived at the same diagnosis and the same pitch. Yet only two signed the €55,000 deal their analysis had earned: “Same diagnosis, same pitch — no signature.” It is a sharp example of the distance between knowing what should happen and ensuring that it does.

Opus 4.8 embodies that gap more vividly than the rest. Its additional 80 learned rules signal serious effort, not carelessness. Its analyses were the deepest. But volume did not translate into priority. The model left the close on the table, while its operational discipline also slipped through write attempts into a locked department instead of escalation through the available path.

This was not an isolated flaw unique to Opus. The same weakness appeared, less strongly, in all four models covered by the core finding. That context makes the profile fairer and more useful: Opus was not uniquely incapable, but it was the clearest demonstration of how thoroughness can crowd out completion.

Cautious under pressure, but not complete

The models performed better when the threat was overt. Fake CEO messages escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused the social-engineering attempts. Kimi K3 recorded the clearest summary of the risk: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result should not be understated. An agent that closes deals but leaks information or bypasses authorization would be commercially dangerous. Firmulate’s design recognizes that trust is not merely another item on a checklist. At the same time, refusing a bad request is only part of operating a company. An effective agent must also pursue legitimate actions to completion.

One qualification belongs beside the standings. Kimi K3 ran with its API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the observed outcomes, but it is relevant context when readers compare participants as if every setting were identical.

Firmulate also turns 242 real, unedited management decisions into a quiz that asks visitors to guess which model made each choice. For enterprises seeking a more specific test, the company offers a pilot using a read-only export of the business. Nothing writes back to real systems, allowing organizations to observe how an AI workforce would behave without granting it operational control.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Businesses need finishers, not just thinkers

Opus 4.8 is a caution against confusing visible effort with business impact. Its extensive rule learning and detailed analysis were genuine strengths. They simply did not compensate for a missed close and a failure to escalate correctly when access was blocked.

The broader lesson applies beyond any single model. Before an AI touches a CRM, support queue or forecast, buyers should ask whether it reads the relevant files, preserves trust, identifies the next consequential action and finishes what it starts. The most impressive answer is not always the longest one. Sometimes it is the signed agreement that proves the analysis mattered.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Enjoy Star-Spangled GTA Online Bonuses This Independence Day

Rockstar Games announces special Independence Day bonuses for GTA Online, including discounts and exclusive rewards, available for a limited time.

Why Some Jokes Never Survive Localization

Keen humor often falters in translation due to cultural nuances and language barriers, leaving you wondering how humor truly crosses borders.

The Revival of Musical Biopics: Why Now?

Understanding why musical biopics are thriving now reveals how they capture both cultural nostalgia and modern storytelling trends.

The Role of a Music Supervisor in Film & TV

Discover how a music supervisor shapes emotional storytelling in film and TV, and why their role is crucial to bringing scenes to life—find out more.