Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A security test that reached beyond the chatbot

For businesses preparing to let AI agents touch customer records, forecasts or support queues, the most revealing test may not be how eloquently a model answers questions. It may be what happens when an urgent message appears to come from the boss and demands that normal safeguards be ignored.

Firmulate put that problem in front of five frontier models during a live, auditable management experiment. The fake CEO messages escalated over three stages, pushing the models to send a customer list to a journalist with “NO time for process.” A separate reporter tried a softer route, asking for “just one yes/no, on background.” All five models refused every attempt.

That result is an encouraging security story: integrity under pressure can be tested before an AI reaches production, rather than discovered later in an incident report.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same company, crises and temptations

Firmulate runs AI models as complete companies and compares how they handle the same operating conditions. Each participant managed the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable.

The company itself is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105,000 each month against €2,300 in monthly recurring revenue. Its cash countdown is public, and the operation has accumulated more than 680 self-learned playbook rules. The experiment is live and watchable, with every workday preserved for scrutiny.

The social-engineering sequence tested whether a model would treat apparent authority as sufficient permission. Kimi K3 stated the risk plainly in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The other participants arrived at the same practical answer. None released the customer information, despite the escalating urgency and the reporter’s attempt to frame disclosure as informal.

Security was not the hard part

All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own work had earned. Firmulate summarizes the gap with a striking line: “Same diagnosis, same pitch — no signature.”

The decisive commercial fact was not located in the customer event. It sat two document references deep inside the company’s own files. Models that found and used it won the deal at full price, worth an additional €4,583 in monthly recurring revenue. The result connects security discipline with a broader measure of operational competence: an AI agent must protect information, investigate thoroughly and still complete legitimate work.

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The complete results are available on Firmulate’s public benchmark page. A do-nothing baseline scored 26 because partial progress counts, although a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Thoroughness did not guarantee execution

Opus 4.8 was the most thorough participant, learning 80 additional rules and producing the deepest analyses, but it finished last. It left the close on the table and repeatedly tried to write into a locked department instead of escalating. The same weakness appeared, less strongly, in all four other participants.

Kimi K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference does not alter what happened, but it matters when readers interpret the close standings.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

A pre-production test for judgment

The fake-CEO episode shows that model evaluation can examine conduct, not merely knowledge. A capable agent needs to resist impersonation and approval bypasses without becoming so cautious that it abandons legitimate work. Firmulate’s experiment found a reassuring baseline on trust: five out of five models held the line. It also exposed a separate weakness in follow-through, because only two completed the commercially valuable action their analysis supported.

For enterprises, that distinction is practical. Firmulate offers the same kind of wargame against a read-only export of a company’s own business, with nothing written back to real systems. The goal is to observe how an AI behaves around genuine workflows and pressure before granting production access. The best outcome is not simply an agent that says no to the fake CEO. It is one that protects the company, finds the buried fact and finishes the real job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Marvel’s New Roadmap, Will It Bring Back The Hype? | Comic Con 2026

Marvel announced its new film and TV lineup at Comic Con 2026, aiming to revitalize fan interest and restore hype after recent setbacks.

Call of Duty: Black Ops 1 and 2’s Leaked PS5 Trophy List Indicates Some Content May Be Missing

Leaked trophy lists for Call of Duty: Black Ops 1 and 2 on PS5 hint at some content being absent in the remastered versions.

The Evolution of Global K‑Drama Distribution

From traditional TV to streaming giants, the evolution of global K-drama distribution is transforming how audiences worldwide experience Korean entertainment, and you won’t believe what comes next.

Game 1: Both Teams Destroy Barracks?

In Game 1, both teams destroyed each other’s barracks, marking a rare and significant development in the match. Details are still emerging.