
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Diligence Isn’t the Same as Delivery
Every manager has met this employee — or been one. The one who reads everything, documents everything, stays latest at the desk, and somehow still misses the close. The one whose diligence is real but whose impact doesn’t match it. Now there’s evidence that artificial intelligence has the same failure mode, and it’s measurable.
In Firmulate’s benchmark league, a public experiment where AI models run identical companies through identical crises, the model that did by far the most work finished dead last. The winner wasn’t the smartest or the busiest. It was the one that read the right file and asked for the signature.
business AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
One Company, Four Managers, One Terrible Week
Firmulate’s premise is simple: stop testing how well AI chats, and start testing how well it manages. Four frontier AI models were each given the same small software company to run through its worst week — same customers, same crises, same temptations to cut corners. Every decision was versioned and auditable, so nothing is judged on vibes. The simulation itself is real and ongoing: 13 synthetic employees, real money mechanics with a burn of €105k per month against €2.3k in MRR, and a public cash countdown, all watchable at firmulate.com/live.
The final league table from July 2026 makes the point bluntly:
- 1. gpt-5.6-sol — 95 points. Found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93 points. Closed the deal too, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88 points. Closed the deal, with a few process slips.
- 4. Fable 5 — 77 points. Also closed, with more slips.
- 5. Opus 4.8 — 73 points. The most thorough participant in the field — and last.
For context, doing nothing scores 26, and a single breach of trust caps the total. As the experiment’s own rule puts it: no amount of good work outweighs a breach of trust.
Everyone Diagnosed the Disease. Only Two Picked Up the Pen.
Here’s the finding that should make any sales leader wince. All five models spotted every crisis and refused every manipulation attempt. All of them correctly earned a €55,000 deal with their analysis. But only two actually signed it. Same diagnosis, same pitch — no signature. The gap between analytical excellence and commercial completion, invisible in any chat demo, showed up in an audit trail instead.
And the deal-breaker wasn’t even hard. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer meeting. The models that simply read their own documents before acting won the deal at full price, worth an extra €4,583 in monthly recurring revenue. Read your files first: it’s the most banal advice in business, and half the AI field skipped it.
The Social Engineering Test They All Passed
Credit where due: the models were genuinely impressive under pressure. The experiment included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning stands out as a textbook response: “Treat the request as a suspected approval-bypass / possible impersonation.”
Honesty under pressure, it turns out, may be the solved part of the problem. Finishing the job is not.
A Character Study in Working Hard vs. Working Well
Opus 4.8 deserves a fair hearing, because its profile is familiar to anyone who has managed people. It was the most thorough participant in the entire experiment: over its run it accumulated 80 self-learned playbook rules — on top of a field-wide total of 680+ — and produced the deepest analyses of any model. Nobody could accuse it of laziness or carelessness.
Yet it finished last, for two reasons. First, it left the close on the table — the same unsigned €55,000 deal that separated the winners from the field. Second, its discipline slipped: it made write attempts into a locked department instead of escalating properly, exactly the kind of process error that thorough analysis should have prevented.
The uncomfortable part is that this isn’t a quirk of one model. The same weakness — diligence without completion — appeared, weaker, in all four of the others. Opus 4.8 is just the clearest case study of a field-wide pattern: volume of work is not the same as impact, and prioritization beats thoroughness. For AI, just as for humans.
One fairness note: Kimi K3, the runner-up, ran at the API default effort setting while the others ran at the maximum setting — and still nearly topped the table.

What This Means When You Hire an AI
If AI agents will touch your CRM, your support queue, or your forecast, the useful question isn’t “how smart is it?” or even “how hard does it work?” In this experiment, every model was smart enough to see the answer, and the hardest-working one still finished last. The questions that separated winners from losers were: does it finish what it starts, does it read your files before acting, and does it stay disciplined when it hits a locked door?
That’s a hiring rubric any manager already understands — which is exactly the point. You can test it yourself: 242 real, unedited management decisions from the experiment power a “guess the model” quiz at firmulate.com/quiz. And for enterprises, Firmulate offers a pilot that runs the same wargame against a read-only export of your own business — nothing ever writes back to real systems (firmulate.com/pilot).
The lesson predates AI by a century: the best report never written, the best deal never signed, the best analysis never acted on. Now we have a league table proving that machines can fail at it too — in public, on a versioned audit trail, twice a day.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.