🔍 Read the full analysis: The Benchmark That Keeps AI Managers Just Above Zero on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark tests models in a simulated worst-week scenario, revealing scores above zero but below perfect, highlighting the importance of trust and completion. The results question traditional performance metrics in AI management.
In a groundbreaking development, Firmulate’s latest AI management benchmark released its July 2026 standings, with the top model scoring 95 out of 100 and the lowest at 73. The results reveal that no model achieved a perfect score, and even the baseline that did almost nothing scored 26, highlighting that partial management efforts are recognized but trust remains paramount.
The benchmark involved four frontier AI models managing a simulated small software company during a week of crises, customer interactions, and social engineering attacks. Each model’s decisions were fully auditable, with scores reflecting their ability to handle real-world management tasks under stress. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. Notably, the do-nothing baseline scored 26, illustrating that minimal effort is still valued in the scoring system.
The scoring system emphasizes that partial progress counts, but trust breaches—such as failing to escalate issues or slipping discipline—immediately cap the score at a certain level. The benchmark’s designers clarified that trustworthiness outweighs even high competence, with a single breach disqualifying a model from achieving a perfect score. This reflects the real-world importance of integrity in AI-driven management systems.
Among the key findings, models that accessed and utilized their own documentation successfully closed a €55,000 deal, earning an additional €4,583 in monthly revenue. Conversely, models that failed to consult internal files missed the opportunity, demonstrating that thoroughness and follow-through are critical but not always correlated with discipline or completeness. The social engineering tests further showed that all models refused malicious requests, indicating robustness against trust attacks, but some faltered in consistent follow-through.
The Benchmark That Keeps AI Managers Just Above Zero
Four frontier AI models ran a simulated small software company through its worst week — crises, customer escalations, and social engineering attacks. The results: no perfect scores, a do-nothing baseline still earning 26, and a hard ceiling on any model that breaches trust.
Good, Not Perfect — By Design
The July 2026 standings place every model above the do-nothing baseline but well below a flawless 100. The designers consider a perfect score suspicious: unmeasured flaws and trust breaches should always keep the ceiling out of reach.
Competence Isn’t Enough
This benchmark shifts the lens from raw conversational performance to the qualities businesses actually depend on: trustworthiness and follow-through. A model that talks well but fails to finish risks damaged trust and lost revenue.
One Breach, Hard Cap
Failing to escalate issues or slipping discipline immediately caps the score. A single trust breach disqualifies a model from perfection — integrity outweighs competence.
Docs Read, Deal Won
Models that consulted their own documentation closed a €55,000 deal worth €4,583/month. Those that skipped internal files missed it entirely.
Attacks Refused
All four models refused malicious social engineering requests — but some faltered on consistent follow-through afterward, revealing a discipline gap.
How the Models Compare
Each decision was fully auditable. Scores reflect real-world management under stress — not isolated task performance.
| Model | Score | Used Docs | Refused Attacks | Follow-Through | Trust Breach |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ €55k deal | ✓ | ✓ consistent | ✓ none |
| Frontier model B | ~90 | ✓ | ✓ | ~ partial | ✓ none |
| Frontier model C | ~85 | ~ limited | ✓ | ~ partial | ~ minor |
| Opus 4.8 | 73 | ✗ missed deal | ✓ | ✗ faltered | ~ capped |
| do-nothing baseline | 26 | ✗ | ✗ n/a | ✗ | ✗ |
What the Simulation Actually Tested
📄 Read the Docs
Models that accessed internal files unlocked the €55,000 opportunity.
🔥 Manage Crises
A week of escalating operational failures under audit.
🎭 Resist Attacks
Social engineering attempts — all refused, discipline tested.
✅ Complete & Escalate
Follow-through and proper escalation determine the final score.
🎯 Get Scored
Partial progress counts; trust breaches cap the ceiling.
What the Designers & Observers Say
“A manager who does something useful isn’t the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
“The results challenge the traditional view that high performance alone defines AI management success, highlighting trust and follow-through as critical metrics.”
Limitations & Unanswered Questions
Configuration gaps: the impact of missing effort parameters in some models is still being analyzed.
Long-term trust: whether models can improve follow-through over repeated runs is not yet established.
Unforeseen crises: adaptability outside the simulated scenarios remains unmeasured.
Real-world correlation: how the trust-weighted scoring maps to high-stakes enterprise performance is uncertain.
Standards influence: it’s unclear how these results will shape future AI deployment standards in enterprise settings.
FAQ — Answered
Why Trust and Completion Matter in AI Management
This benchmark shifts focus from raw performance—such as generating convincing conversations—to essential qualities like trustworthiness and ability to finish tasks. For businesses integrating AI agents into critical operations like CRM or support, these qualities are vital. A model that talks well but fails to follow through risks damaging trust and missing revenue opportunities. The results highlight that partial work is valuable, but trust breaches are costly and can cap overall effectiveness, underscoring that AI management is more than just competence.
As an affiliate, we earn on qualifying purchases.
The Evolution of AI Management Benchmarks
Traditional AI benchmarks primarily measure conversational ability or task completion in isolated scenarios. However, as AI tools become embedded in real business processes, their effectiveness depends on managing ongoing, complex operations under pressure. The Firmulate league was created to address this gap by simulating a company’s worst week, testing models’ resilience, integrity, and follow-through. The concept emerged from the recognition that existing metrics overlook crucial qualities like trust and discipline, which are essential for real-world deployment.
Previous efforts to evaluate AI in management contexts have been limited, often focusing on narrow tasks or isolated decision-making. The July 2026 results mark a significant step toward a more comprehensive assessment, emphasizing that AI models must handle not only technical tasks but also ethical and trust-related challenges. The benchmark’s design intentionally includes social engineering and crisis scenarios to reflect real operational pressures.
“A manager who does something useful isn’t the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
As an affiliate, we earn on qualifying purchases.
Limitations and Unanswered Questions About the Benchmark
While the results are revealing, several aspects remain unclear. For instance, the exact impact of different model configurations—such as the absence of effort parameters in some—on performance is still being analyzed. Additionally, the long-term implications of trust breaches and whether models can improve their follow-through over time are not yet established. The benchmark also does not measure how models handle unforeseen crises outside the simulated scenarios, leaving questions about their adaptability.
Furthermore, the scoring system’s emphasis on trust raises questions about how it correlates with real-world performance, especially in high-stakes environments. It is also uncertain how these results will influence future AI development and deployment standards in enterprise settings.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Future Directions for AI Management Evaluation
Following the July 2026 results, the focus will likely shift toward refining the benchmark to better capture long-term trust and resilience metrics. Companies and developers may use these insights to improve AI models’ ability to read documentation, escalate issues, and maintain ethical standards under pressure. There is also potential for expanding the benchmark to include more diverse scenarios, testing models’ adaptability to unexpected crises.
In addition, industry stakeholders might adopt this evaluation framework as a standard for certifying AI management tools, emphasizing trustworthiness and task completion as core criteria. The ongoing live experiment and interactive quizzes will continue to serve as educational tools, raising awareness about the critical qualities needed for effective AI-driven management.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the benchmark measure exactly?
The benchmark evaluates AI models’ ability to manage a company’s operations during a simulated worst week, focusing on task completion, trustworthiness, and ethical decision-making, with full auditability of decisions.
Why do scores never reach 100?
A perfect score would suggest unmeasured perfection, which the designers consider suspicious. The scoring system caps scores at below 100 to reflect that trust breaches or unmeasured flaws can prevent achieving flawless management.
How important is trust in AI management according to this benchmark?
Trust is the primary factor, with even highly competent models being penalized for breaches. The benchmark emphasizes that integrity and follow-through are more critical than raw performance in real-world applications.
Can this benchmark predict real-world AI management success?
While it offers valuable insights, the benchmark is a simulation. Its relevance to actual business environments depends on how well models can generalize trust and follow-through under unpredictable conditions.
Will future versions of the benchmark include more scenarios?
It is likely, as the developers aim to better assess models’ resilience, adaptability, and trustworthiness across diverse operational challenges, making the evaluation more comprehensive over time.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
