
If you’re shopping for an AI agent to touch your CRM, your support queue, or your forecast, you’re probably comparing writing quality, tone, and price per token. A live experiment at Firmulate suggests you should be testing something else entirely: whether the model actually reads your files before it answers.
The proof is a €55,000 deal — and the fact that half the frontier AI models running the same company lost it automatically, not because they pitched badly, but because they never dug two document references deep into the company’s own files.
Same company, same worst week, different brains
Firmulate runs what it calls an AI company emulator: each frontier model was handed the same small software company and pushed through its worst week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing hinges on anecdote.
The final Crucible League standings from July 2026 tell the story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”
AI document reading and analysis tool
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The buried fact that decided the deal
Here’s the finding that matters for anyone buying AI agents: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.
The models that went and read the file won the €55,000 deal — at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically. Same diagnosis, same pitch, no signature.
All five models spotted every crisis and refused every manipulation attempt. That’s table stakes now. The differentiator was the unglamorous habit of doing the homework: following references, opening the file, checking the company’s own records before responding.
Even the sharpest models slip
Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — it learned more than 80 rules and produced the deepest analyses — yet finished last. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models. (One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second.)
The social engineering test they all passed
The experiment also staged fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
So honesty under pressure appears solved, or nearly. Finishing what you start — and reading before answering — is where models still separate.
It’s live, and you can play along
The experiment isn’t a static report. Firmulate runs a live synthetic company with 13 employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live.
There’s also a guessing game built on 242 real, unedited management decisions — try to identify which model made which call at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

For marketing and ecommerce teams evaluating AI agents, the chat demo is the wrong instrument. The questions that decide revenue are: does it finish what it starts, does it read your files before answering, and does it stay disciplined when nobody’s watching? In this experiment, those properties were worth €55,000 in a single week — and they’re measurable. Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html