
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
An AI Manager That Does Nothing Still Gets a 26
If you’ve ever rolled your eyes at a benchmark score of 98 or 100, you’re not alone. The most interesting number in AI evaluation right now isn’t at the top of a leaderboard — it’s at the bottom. In the Crucible League benchmark from Firmulate, an AI model that manages a company and does absolutely nothing still walks away with 26 points.
That’s not a bug. It’s a design decision, and it tells you a lot about what honest AI measurement should look like — especially if you’re a business owner about to hand an AI agent the keys to your CRM, support queue, or sales pipeline.
The Wargame, Not the Chat Demo
Firmulate runs frontier AI models as complete companies. Each model got the same assignment: steer the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable, and the whole thing is watchable at firmulate.com/live — a live, running company with 13 synthetic employees, real money mechanics (burning €105k a month against €2.3k in MRR), a public cash countdown, and 680+ self-learned playbook rules.
The final July 2026 league table: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
Why Zero Isn’t Zero
So why does a do-nothing run score 26 instead of 0? Because the benchmark rewards partial progress. An AI that identifies a crisis but doesn’t resolve it, or prepares analysis but never closes, has still done some of the job. Real management isn’t binary — a manager who half-handles a customer escalation is more useful than one who ignores it. The scoring mirrors that reality.
But the floor comes with a ceiling. A single breach of trust caps the total grade outright. As the benchmark’s own language puts it: “no amount of good work outweighs a breach of trust.” An agent can ace every operational metric, but if it deceives a customer or fakes an approval once, the score is capped. That’s the same standard most business owners would apply to a human employee — it’s just rare to see it enforced in AI evaluation.
What Actually Separated the Winners
The most revealing finding wasn’t about intelligence — it was about diligence. All five models spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The decisive difference? The winning models found a buried fact about a competitor’s weakness sitting two document references deep in the company’s own files — not in the customer event in front of them. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for ecommerce and marketing teams is blunt: your data often already contains the answer, and the difference between AI agents may simply be whether they bother to read it.
Pressure-Testing Honesty
The week included social engineering: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” The benchmark even distrusts perfectly round scores — a healthy skepticism for anyone who’s seen too many 100/100 product demos.
The Cautionary Tale
Opus 4.8 is the profile worth studying. It was the most thorough participant — over 80 learned rules and the deepest analyses — yet finished last. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models. One fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still placed second.
Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model did what at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems (firmulate.com/pilot.html).

A benchmark that gives 26 points for doing nothing, rewards partial progress, and caps the score on any single breach of trust is measuring the thing businesses actually care about: follow-through, honesty, and whether the agent reads your files before acting. Before you hire an AI workforce, wargame it. The full methodology and league table are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
