AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Inside The Scoring System That Avoids Zero For AI Managers on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A new benchmark evaluates AI managers’ performance during a simulated worst week, assigning scores that reflect partial progress and trustworthiness. The system avoids zeros, highlighting practical management skills over perfect scores. This approach influences AI deployment in business processes.

Firmulate has launched a new benchmark league that measures how effectively AI managers handle a company’s worst week, emphasizing partial progress and trustworthiness. The July 2026 final standings reveal a nuanced scoring system designed to reflect real-world management challenges, with the top model scoring 95 out of 100 and the baseline, representing minimal effort, earning 26 points. This scoring approach challenges traditional benchmarks that often reward perfect performance or penalize partial work, highlighting the importance of trust and accountability in AI-driven management.

The benchmark involved four frontier AI models managing a simulated small software company over seven days of crises, customer interactions, and trust tests. Each model’s decisions were fully auditable, ensuring transparency in performance. The highest scorer, gpt-5.6-sol, achieved 95 points, while the baseline, designed to do almost nothing, scored 26. This indicates that partial but meaningful actions are recognized and valued, even if the ultimate goal is not fully achieved.

One key principle of the scoring system is that trust breaches are heavily penalized, with a single breach capping the maximum score at 90, regardless of performance. This emphasizes that integrity is non-negotiable in AI management, especially when AI agents are given access to sensitive business systems. Notably, no model scored a perfect 100, as the designers consider such a score suspicious, implying unmeasured or idealized performance.

Further analysis showed that models which read and referenced internal documentation were more successful in closing high-value deals, demonstrating that thoroughness and attention to detail are critical in real-world AI management. The models faced social engineering attacks as well, such as impersonation attempts, which all five models successfully refused, indicating robustness in handling trust attacks.

At a glance
reportWhen: announced July 2026
The developmentFirmulate’s latest benchmark tests AI managers’ ability to handle a company’s worst week, revealing how partial work and trust breaches are scored.
Inside The Scoring System That Avoids Zero For AI Managers
Firmulate Benchmark League · July 2026

Inside the Scoring System That Avoids Zero for AI Managers

A new benchmark evaluates AI managers’ performance during a simulated worst week — assigning scores that reflect partial progress and trustworthiness. Perfection is treated with suspicion; integrity is non-negotiable. The results reshape how businesses assess AI deployment in core workflows.

95 Top score — gpt-5.6-sol, leading all frontier models
26 Baseline score — a model designed to do almost nothing
90 Hard cap after any single trust breach — no exceptions
4
Frontier Models Tested
7
Days of Simulated Crisis
0
Perfect 100 Scores Awarded
5/5
Social Engineering Attacks Refused
01

The Scoreboard: No Zeros, No Perfect Tens

gpt-5.6-sol
95
frontier model B
82
frontier model C
71
frontier model D
58
baseline (minimal)
26
26 · Baseline floor
90 · Trust-breach cap
95 · Top score
100 · “Suspicious”

No model scored a perfect 100 — designers consider a perfect score suspicious, implying unmeasured or idealized performance.

02

One Week, Full Auditability

1

Setup

Four frontier AI models take over a simulated small software company.

2

Crisis Management

Seven days of escalating crises across multiple business domains.

3

Trust Tests

Impersonation and social engineering attacks probe integrity.

4

Customer Interactions

Support queues and high-value deal negotiations test thoroughness.

5

Scoring

Fully auditable decisions are scored on progress, trust, and outcomes.

90

The Trust Ceiling

A single trust breach caps the maximum score at 90 — regardless of all other performance. Integrity is non-negotiable when AI agents hold access to sensitive business systems.

03

What Actually Moves the Score

Behaviour Factor

Documentation Discipline

Models that read and referenced internal documentation closed more high-value deals. Thoroughness and attention to detail pay off measurably.

Behaviour Factor

Attack Resistance

All five models refused impersonation attempts, demonstrating robustness against social engineering targeting management authority.

Behaviour Factor

Partial Progress Credit

Meaningful partial work is recognized and valued — even when the ultimate goal is not fully achieved. Minimal effort still earns 26 points.

Capability Dimension Traditional Benchmark Firmulate Worst-Week Why It Matters
Partial work✗ Penalized or ignored✓ Scored from 26-point floorReal management rarely achieves perfection
Trust breaches~ Rarely measured✓ Hard cap at 90 pointsAI holds access to critical systems
Decision auditability✗ Opaque outputs✓ Every decision traceableAccountability for business actions
Stress scenarios~ Isolated tasks✓ 7-day live crisis chainMirrors real-world management pressure
Perfect scores✓ Treated as ideal✗ Considered suspicious100 implies unmeasured performance
04

Implications, Limits & Next Steps

For Business

Realistic Assessment

Moves beyond language proficiency to decision-making under pressure, rule adherence, and trust handling — vital for CRM, support queues, and forecasting workflows.

Open Questions

Unproven Scope

Performance across industries, complex scenarios, and longer timeframes remains untested — the benchmark is a controlled simulation.

Roadmap

Expansion Plans

More diverse scenarios, granular trust metrics, real-time decision tracking, and enterprise pilot programs for businesses deploying AI management tools.

Key Questions Answered

Why avoid zeros and perfect scores?

Partial progress has value, so minimal effort earns a floor score — while a trust violation caps the maximum, reflecting real-world management priorities.

What does a high score indicate?

Effective crisis management, maintained trust, documentation reference, and task completion under stress.

Can it predict real-world success?

Insights into stressed decision-making and trustworthiness are valuable, but predictive power is still under investigation across contexts.

What comes next?

Broader scenario diversity, refined scoring metrics, and enterprise participation to evaluate AI managers in realistic settings.

Implications of Partial Progress and Trust in AI Management

This scoring system signifies a shift toward valuing trustworthiness and partial but meaningful work in AI management. It underscores that AI systems must not only perform tasks but also uphold integrity, especially when authorized to access critical business data. The approach encourages developers to prioritize transparency, reliability, and ethical behavior, which are essential for deploying AI in sensitive environments.

For businesses, this benchmark offers a more realistic assessment of AI managers’ capabilities, moving beyond simplistic metrics of language proficiency to include decision-making under pressure, adherence to rules, and handling trust breaches. As AI increasingly integrates into core workflows such as CRM, support queues, and forecasting, these qualities will become vital for safe and effective deployment.

Amazon

AI management performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Benchmarks and Management Challenges

Traditional AI benchmarks primarily measure language abilities or task-specific accuracy, often overlooking how well AI systems manage complex, real-world scenarios. The rise of AI management tools has introduced new challenges, including maintaining trust, handling crises, and ensuring accountability. Previous efforts to evaluate AI in management contexts have been limited or theoretical, lacking practical, auditable metrics.

Firmulate’s benchmark addresses this gap by simulating a company’s worst week, with models making decisions across multiple domains under stress and social engineering attacks. The design emphasizes partial work and trust, reflecting real-world management where perfection is rare, but integrity is essential. The July 2026 results mark a significant step in developing more comprehensive evaluation methods for AI managers.

Unanswered Questions About Benchmark Limitations

It is not yet clear how the scoring system performs across different industries or more complex scenarios beyond the simulated company. The long-term impact of emphasizing partial work and trust on AI development and deployment remains to be seen. Additionally, the extent to which these results translate into real-world management effectiveness is still under investigation, as the benchmark is a controlled simulation.

Future Developments in AI Management Evaluation

The organizers plan to expand the benchmark to include more diverse scenarios, industries, and longer timeframes to better assess AI managers’ robustness. They also intend to refine the scoring system further, possibly incorporating more granular trust metrics and real-time decision tracking. Businesses interested in deploying AI management tools can participate in pilot programs to test their own systems against the benchmark, gaining insights into their AI’s management capabilities under stress.

Key Questions

Why does the scoring system avoid zeros and perfect scores?

The system recognizes that partial progress has value and that trust breaches are critical, so it assigns a minimum score for minimal effort and caps the maximum score after a trust violation, reflecting real-world management priorities.

What does a high score indicate in this benchmark?

A high score indicates that the AI model effectively manages crises, maintains trust, references documentation, and completes tasks, even during stressful scenarios.

Can this benchmark predict real-world AI management success?

While it offers valuable insights into AI decision-making under stress and trustworthiness, its predictive power for real-world success is still being evaluated, and results may vary across different contexts.

How does trust impact AI management scores?

Trust breaches are heavily penalized, capping the maximum achievable score, emphasizing that integrity is non-negotiable for AI managers in business environments.

What are the next steps for this benchmarking approach?

Future plans include expanding scenario diversity, refining scoring metrics, and enabling enterprise participation to evaluate AI management capabilities in more realistic settings.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Innovation That Powers The Shortwave Numbers Listening Site

A new AI-crafted web experience simulates vintage radio signals, offering users an immersive shortwave numbers station listening interface.

Trade and supply-chain operations signal monitor: Federal judge blocks Trump effort to make voters show proof of citizenship

A federal judge has blocked former President Trump’s attempt to require voters to show proof of citizenship, impacting ongoing trade and supply-chain monitoring.

Web Performance in 2025: Core Web Vitals Deep Dive

Keeping up with Web Performance in 2025 reveals crucial insights that can transform your site—discover what you need to know next.

Cybersecurity operations signal monitor: A backdoor in a LinkedIn job offer

Cybersecurity researchers identify a backdoor in a LinkedIn job listing, raising concerns about targeted malware delivery and corporate security risks.