📊 Full opportunity report: The AI Leaderboard That Reveals True Capabilities Post-Demo on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live benchmark by Firmulate tests AI models as managers in a simulated company crisis. Results show models excel at diagnosis but often fail in execution and trust, revealing a gap in true management skills. This development shifts focus toward evaluating AI for operational decision-making, not just chat quality.
Firmulate has launched a live management benchmark experiment that tests AI models as managers in a simulated company facing its worst week. The results, announced after the July 2026 Crucible League, show that while models can identify crises accurately, they often fail to complete critical tasks like signing deals or escalating issues, highlighting a gap in true management skills. This matters because it shifts the focus from chat and coding benchmarks to real-world operational performance, which is crucial for AI adoption in business decision-making.
The experiment involved five AI models managing a small software company with real money mechanics, as detailed in the original analysis, burning €105,000 monthly against €2.3 million in revenue. The models were evaluated on their ability to diagnose crises, communicate effectively, make decisions, and uphold trust in scenarios such as negotiations, crisis escalation, and manipulation attempts. The top performer, gpt-5.6-sol, scored 95 out of 100, while others lagged behind, with Opus 4.8 scoring just 73. Despite strong diagnostic skills, only two models signed the €55,000 deal their own analysis identified as the best opportunity, illustrating a critical execution gap. The experiment also tested models’ resistance to social engineering, with all five refusing fake CEO requests, demonstrating competence in safeguarding against manipulation.
However, the models that produced the most detailed analyses often failed in the final steps—signing deals or escalating issues—highlighting that thoroughness does not equate to effective management. For example, Opus 4.8, which added 80 rules and produced deep analyses, finished last because it failed to escalate a problem into the right department. Contextual factors such as different API effort parameters were noted but did not change the overall ranking significance. The experiment’s design enforced a strict trust standard: a single breach capped the score, emphasizing the importance of honesty over mere performance.
Implications for AI in Business Management
The results demonstrate that current AI models can diagnose problems and resist manipulation but often struggle with completing management tasks that require judgment, escalation, and trust. This suggests that AI evaluation should extend beyond chat quality and coding benchmarks to include operational effectiveness in real-world scenarios. For organizations considering AI for decision-making, the key takeaway is the importance of testing models’ ability to read organizational context, prioritize correctly, and uphold trust over time. The experiment underscores that effective management involves more than generating detailed reports; it requires execution, escalation, and maintaining integrity under pressure.
This shift in evaluation focus could influence how AI tools are integrated into enterprise workflows, emphasizing the need for rigorous testing of management capabilities rather than superficial performance metrics. It also highlights that AI models might need additional training and safeguards before being entrusted with critical operational decisions, especially in high-stakes environments.
As an affiliate, we earn on qualifying purchases.
Background on AI Benchmarking and Management Testing
Traditional AI benchmarks primarily assess models based on their ability to generate accurate responses, code, or engage in human-like conversation. These metrics, however, do not capture how models perform in complex, real-world management scenarios where trust, escalation, and decision-making under pressure are critical. Firmulate’s experiment is a response to this gap, creating a live, dynamic environment where models manage a simulated company facing crises, with real financial stakes and decision logs. The July 2026 Crucible League marked the culmination of this effort, providing a comparative ranking of models based on their management performance.
Previous benchmarks have focused on technical prowess, but recent developments suggest that operational competence—such as reading organizational files, resisting manipulation, and completing tasks—is equally vital. The experiment builds on this understanding, emphasizing that true AI management capability involves consistent, trustworthy execution across diverse scenarios, not just diagnostic accuracy.
“This experiment reveals that diagnosing a crisis is only half the job. Effective management requires trust, execution, and escalation—areas where current models still fall short.”
— Thorsten Meyer, Lead Developer at Firmulate
Unresolved Questions About AI Management Performance
It remains unclear how well these models will perform in different industries or with more complex organizational structures. The experiment was conducted in a controlled, simulated environment with a small company, which may not fully capture the nuances of larger, real-world enterprises. Additionally, the long-term reliability of models in maintaining trust and executing management tasks over extended periods has not yet been established. Further testing across diverse scenarios and industries is necessary to confirm whether these findings generalize broadly.
Next Steps for AI Management Benchmarking
Following the July 2026 results, firms like Firmulate plan to refine their benchmarks by incorporating more complex scenarios, longer management cycles, and larger organizational contexts. There is also interest in developing evaluation metrics that better capture trustworthiness, escalation quality, and decision effectiveness. Researchers and enterprises will likely explore hybrid approaches combining AI models with human oversight to mitigate execution gaps. Additionally, organizations considering AI for management roles should start testing models in simulated environments tailored to their specific operations before deployment.
Further public benchmarks and live testing platforms are expected to emerge, fostering a more comprehensive understanding of AI’s operational capabilities and limitations in management roles.
Key Questions
What is the main purpose of Firmulate’s management benchmark?
The benchmark aims to evaluate AI models’ ability to manage a company effectively, including diagnosing crises, making decisions, escalating issues, and maintaining trust, beyond just generating responses or code.
How do current AI models perform in management tasks according to the experiment?
Models are proficient at diagnosing problems and resisting manipulation but often struggle with completing tasks like signing deals, escalating issues properly, and maintaining consistent trustworthiness over time.
Why is this shift in evaluation important?
It emphasizes operational effectiveness and trustworthiness, which are critical for deploying AI in real-world management roles, moving beyond superficial benchmarks that do not measure practical management skills.
Will this testing approach replace traditional benchmarks?
Not entirely, but it will complement existing metrics by providing a more comprehensive assessment of AI’s operational and management capabilities in realistic scenarios.
What should companies do before deploying AI in management roles?
Organizations should run their own simulated management tests, focusing on trust, escalation, and decision-making, to ensure models can handle their specific operational complexity.
Source: ThorstenMeyerAI.com