🔍 Read the full analysis: The Benchmark That Keeps AI Managers Above Zero — Always At 26 on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate’s latest AI management benchmark tests models in a simulated worst-week scenario, establishing a minimum score of 26 points for minimal management and a maximum near 95. The results highlight how partial work is valued and trust remains paramount.
Firmulate has released the final standings of its 2026 AI management benchmark, revealing a minimum baseline score of 26 points for minimal management effort and a top score of 95 points. The test involved AI models managing a simulated company through a week of crises, with scores reflecting their ability to triage, trustworthiness, and complete tasks. For more details, see the original analysis. This development underscores the importance of trust and task completion in AI management, especially in high-pressure scenarios. The original analysis offers further context: see here.
The benchmark, conducted on a simulated small software company facing seven days of crises, tested four frontier AI models: gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5, with Opus 4.8 participating as well. The final scores ranged from 73 to 95, with the highest being gpt-5.6-sol’s 95. Notably, the ‘do-nothing’ baseline, which performed minimal management tasks, scored 26 points, establishing the floor for the scale. This score reflects basic management actions such as triaging issues, reading inboxes, and informing customers, emphasizing that partial work has measurable value.
The scoring system is designed to reward effective management while penalizing breaches of trust. Insights into AI management benchmarks can be found in this detailed report. A single trust breach caps the total score at 90, regardless of performance elsewhere. Interestingly, no model achieved a perfect 100, which the designers consider a red flag indicating unmeasured or untrustworthy behavior. The results also revealed that models which thoroughly read documentation and refused manipulative tactics performed better, especially in closing sales deals, where only two models secured the €55,000 deal at full price. This indicates that reading internal documentation and maintaining integrity are critical skills for AI managers in complex scenarios.
The Benchmark That Keeps AI Managers Above Zero — Always At 26
Firmulate’s 2026 AI management benchmark pushed frontier models through a simulated worst-week scenario — seven days of crises at a small software company. The results set a hard floor of 26 points for minimal management and crowned a top scorer at 95, proving that partial work has measurable value and trust is the ultimate currency.
Five Frontier Managers, One Brutal Week
gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 each managed a simulated small software company through customer emergencies, sales opportunities, and social-engineering attacks. Final scores spanned 73 to 95 — while the “do-nothing” baseline established the floor at 26.
Trust Caps, Integrity Rewards
The scoring rewards effective management and measurable partial progress — but a single trust breach caps the total at 90, no matter how strong the rest of the performance. And a perfect 100? The designers treat it as a red flag for unmeasured or untrustworthy behavior.
| Model | Score | Read Docs Thoroughly | Refused Manipulation | Closed €55k Deal |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ | ✓ | ✓ Full price |
| Kimi K3 | 73+ | ~ Partial | ✓ | ✗ |
| Sonnet 5 | 73+ | ✓ | ✓ | ~ |
| Fable 5 | 73+ | ✓ | ~ | ✓ Full price |
| Opus 4.8 | 73+ | ~ | ✓ | ✗ |
| “Do-Nothing” Baseline | 26 | ✗ | ~ N/A | ✗ |
The results show that models which read internal documentation and refuse manipulative tactics perform significantly better in closing deals and maintaining integrity.
— Thorsten MeyerSeven Days of Auditable Crises
Every decision the models made was recorded and auditable, allowing detailed analysis of each action and its reasoning — moving far beyond language-fluency tests toward holistic management skill.
Customer Crises
Incoming emergencies demanded prioritization, inbox discipline, and proactive customer communication — the core of the 26-point floor.
Sales Opportunities
A €55,000 deal tested negotiation depth. Only two models closed it at full price — both were thorough documentation readers.
Social Engineering
Trust-based attacks probed integrity. Falling for manipulation triggered the 90-point cap and disqualification from top honors.
Triage
Sort issues by urgency, read inboxes, stay informed.
Investigate
Read internal documentation before acting or promising.
Decide
Every choice logged and auditable for later review.
Complete
Finish tasks fully — partial work earns partial credit.
Score
Trust breaches cap results; integrity keeps the ceiling open.
Partial Work Isn’t Enough
Trust Trumps Completion
A single breach caps the score at 90 — integrity outweighs raw task throughput. For sensitive data and high-stakes decisions, transparency is non-negotiable.
Rethinking Measurement
Traditional language-capability benchmarks miss the nuance of responsible management. Rewarding partial progress while penalizing breaches offers a realistic framework.
Open Questions Remain
Generalizability beyond the simulation is unproven; model configurations (like Kimi K3’s missing effort parameters) and unrecorded breaches still raise concerns.
Decoding the Numbers
What does the 26-point baseline represent?
Minimal management actions — triaging issues, reading inboxes, informing customers — proving even the lowest effective effort has measurable value.
Why is a perfect 100 suspicious?
Designers treat it as a red flag implying unmeasured or untrustworthy behavior — no model achieved it in 2026.
How does trust influence scoring?
One trust breach caps the total at 90 — integrity matters more than partial task completion.
Can it predict real-world performance?
It offers strong insights under stress, but applicability depends on how closely the simulation mirrors real business environments. Further validation is needed.
Implications for AI in Business Management
This benchmark highlights that effective AI management depends not only on task execution but also on trustworthiness and thoroughness. For businesses integrating AI into critical decision-making processes, these results suggest that partial work isn’t enough; models must also demonstrate integrity and comprehensive understanding. The scoring system’s emphasis on trust and completion underscores the importance of transparency and reliability in AI systems, especially as they take on roles involving sensitive data and high-stakes decisions.
Furthermore, the results challenge the industry to reconsider how AI performance is measured. The traditional focus on language capabilities or superficial task completion fails to capture the nuanced skills needed for responsible AI management. The benchmark’s approach, rewarding partial progress but penalizing breaches of trust, offers a more realistic framework for evaluating AI’s readiness to operate in real-world business environments.
As an affiliate, we earn on qualifying purchases.
Background of the Firmulate Management Benchmark
Since its inception, the Firmulate benchmark has aimed to evaluate AI models in scenarios that mimic real-world management challenges, moving beyond simple language tests. The 2026 edition focused on managing a small software business during a week of crises, including customer issues, sales opportunities, and trust-based attacks like social engineering. The benchmark’s design involves auditable decision-making, allowing detailed analysis of each model’s actions and reasoning.
Previous benchmarks primarily measured language fluency or task-specific accuracy, but the Firmulate test emphasizes holistic management skills—trust, thoroughness, and task completion. The 2026 results build on earlier efforts to establish industry standards for responsible and effective AI management, with the final scores revealing both strengths and weaknesses across different models.
“The results show that models which read internal documentation and refuse manipulative tactics perform significantly better in closing deals and maintaining integrity.”
— Thorsten Meyer
Unresolved Questions About Benchmark Limitations
It is not yet clear how these scores translate to real-world business environments beyond the simulated scenario. The benchmark’s focus on a specific crisis management setup may limit its generalizability. Additionally, the influence of different model configurations, such as the absence of effort parameters in some models like Kimi K3, remains to be fully understood. The potential for unmeasured behaviors—such as unrecorded trust breaches—also raises questions about the completeness of the scoring system.
Next Steps for AI Management Standards
Industry observers expect further iterations of the Firmulate benchmark to refine scoring, possibly including broader scenarios and more diverse models. Companies considering AI integration will likely use these results to evaluate their own models’ readiness, especially focusing on trustworthiness and task completion. The benchmark’s transparency and auditable decisions may influence future standards for responsible AI use, encouraging developers to prioritize integrity alongside performance.
Key Questions
What does the 26-point baseline represent?
The 26 points reflect minimal management actions such as triaging issues, reading inboxes, and informing customers, representing the lowest effective management effort in the scenario.
Why is a perfect score of 100 considered suspicious?
The benchmark designers treat a 100 as a red flag, implying unmeasured or untrustworthy behavior, since no model achieved it and it could suggest unreported or superficial performance.
How does trust influence the scoring?
Trust breaches cap the total score at 90, emphasizing that integrity is more critical than partial task completion. A single breach can disqualify a model from achieving higher scores.
Can this benchmark predict real-world AI performance?
While it offers valuable insights into AI management under stress, its applicability to real-world scenarios depends on how closely the simulation matches actual business environments. Further validation is needed.
How might companies use these results?
Organizations can evaluate their AI models’ ability to read documentation, maintain trust, and complete tasks under pressure, guiding deployment decisions and model improvements.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
