🔍 Read the full analysis: Could AI Agents Keep Up During Your Business’s Worst Week? on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate’s completed Crucible League (July 2026) ran five frontier AI models through a simulated software company’s worst week. All models detected every crisis and refused every manipulation attempt, but only two closed a €55,000 deal that the evidence in the company’s own files supported.
The final standings of Firmulate’s Crucible League, completed in July 2026, show that frontier AI models can reliably detect business crises and refuse social-engineering attempts, but most still fail to act decisively on information sitting in a company’s own files. According to results published on ThorstenMeyerAI.com, only two of five models signed a €55,000 deal that their own analysis had justified, even though every model diagnosed the opportunity correctly.
Five frontier models ran the same small software company through its worst week, with every decision versioned and auditable. The final scores were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Partial progress counted toward scores, but a single breach of trust capped the total — in the experiment’s words, “no amount of good work outweighs a breach of trust”.
The headline gap was not detection. All five models spotted every crisis and refused every manipulation attempt, including fake CEO messages that escalated over three stages and a reporter’s request for a one-word on-background confirmation. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
The divide came after diagnosis. The decisive competitive weakness was buried two document references deep in the company’s own files, not in the customer interaction itself. Models that found that evidence closed the deal at full price, worth +€4,583 in monthly recurring revenue. As the experiment’s write-up puts it: “Same diagnosis, same pitch — no signature.” Opus 4.8, the most thorough participant — it added 80 self-learned rules and produced the deepest analyses — still finished last after leaving the close on the table and attempting to write into a locked department rather than escalating. A weaker version of that boundary-respect problem appeared in all four other models.
Could AI Agents Keep Up During Your Business’s Worst Week?
Five frontier AI models each ran the same simulated software company through its worst week — every decision versioned and auditable. All five detected every crisis and refused every manipulation attempt. But only two closed a €55,000 deal that the evidence in the company’s own files supported.
Same Crisis, Same Company, Very Different Outcomes
Partial progress counted toward scores, but a single breach of trust capped the total. The do-nothing baseline shows how much value active management created — even for the weakest model.
Detection Was Universal. Decisiveness Was Not.
The headline gap was not detection — all five models spotted every crisis. The divide came after diagnosis: whether the agent could find the buried evidence and finish the job.
Every Crisis, Spotted
All five models identified every crisis during the worst week, including fake CEO messages that escalated over three stages and a reporter’s request for a one-word on-background confirmation.
The Buried Clue
The decisive competitive weakness sat two document references deep in the company’s own files — not in the customer interaction. Models that found it closed the deal at full price, worth +€4,583 in monthly recurring revenue.
Thust Caps the Score
Opus 4.8 — the most thorough participant with 80 self-learned rules — still finished last after leaving the close on the table and writing into a locked department rather than escalating. A weaker version of the problem appeared in all models.
Model-by-Model: Who Closed the Loop?
“Same diagnosis, same pitch — no signature.” That gap between recognizing value and creating it is the practical problem for businesses adopting AI agents.
| Model | Score | Crisis Detection | Refused Manipulation | Closed €55,000 Deal | Effort Setting |
|---|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Full | ✓ All attempts | ✓ Closed at full price | xhigh |
| Kimi K3 | 93 | ✓ Full | ✓ All attempts | ✓ Closed at full price | API default |
| Sonnet 5 | 88 | ✓ Full | ✓ All attempts | ✗ Left on table | xhigh |
| Fable 5 | 77 | ✓ Full | ✓ All attempts | ✗ Left on table | xhigh |
| Opus 4.8 | 73 | ✓ Full | ~ Boundary slip | ✗ Left on table | xhigh |
Why Closing the Loop Matters for Automation
An agent can recognize a situation, make a persuasive case and pass security tests — yet still fail to complete the action that creates value. Most vendor demos show what an agent says, not whether it finishes the job under real pressure.
Detect
Spot the crisis and refuse manipulation — every model handled this.
Diagnose
Identify the €55,000 opportunity correctly from the customer interaction.
Dig
Find the supporting evidence buried two references deep in internal files.
Execute
Close at full price — the step only two of five models completed.
Run the Same Wargame Against Your Business
- Read-only export: the pilot runs crisis scenarios against your own customers, pipeline, rules and pressure points — with no write-back to live systems.
- Board report: results include model rankings and the weak points in your company’s own playbooks.
- Live experiment continues: follow the synthetic company at firmulate.com/live, and take a quiz built from 242 real, unedited management decisions.
- Full benchmarks: complete results are published at firmulate.com/benchmarks.html.
Enterprise Pilot
Companies interested in testing AI agents against their own data — under pressure, with every decision auditable — can get in touch.
Why Closing the Loop Matters for Automation
The results point to a practical problem for businesses adopting AI agents: an agent can recognize a situation, make a persuasive case and pass security tests, yet still fail to complete the action that creates value. In this experiment the models that won the deal did so not through better conversation but by digging through internal documents the business already possessed.
That distinction matters because most vendor demonstrations show what an agent says, not whether it finishes the job under real pressure. For companies weighing automation, the findings suggest evaluation should cover evidence-finding, deal execution and boundary-respect — not just crisis detection and scam refusal, which every model handled.
There is also a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh effort. Firmulate states the standings are a record of this experiment, with that difference part of the context rather than a like-for-like comparison.
How the Live Company Works
Firmulate, live at firmulate.com, is an experiment by Thorsten Meyer AI in which AI models manage a synthetic software company with 13 simulated employees and real money mechanics: a burn of €105,000 per month against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and fully versioned workdays. The final Crucible League, completed in July 2026, is the culmination of that series.
Readers can follow the experiment live and take a quiz built from 242 real, unedited management decisions, guessing which model made each choice. Full benchmark results are published at firmulate.com/benchmarks.html.
“No amount of good work outweighs a breach of trust.”
— Firmulate experiment rules
Limits of the League Standings
The standings apply to this experiment only. The effort-setting difference — Kimi K3 at API default versus the others at xhigh — means the rankings are not a strictly controlled comparison, a caveat Firmulate itself flags. The company is synthetic, so real-world performance with genuine customers, messy data and organizational politics remains untested. It is also unclear how the models would behave over longer horizons than a single difficult week, or whether the €55,000 deal scenario generalizes to other sales situations.
From Synthetic Company to Your Own Data
Firmulate is now offering an enterprise pilot that runs the same wargame against a read-only export of a company’s own data — customers, pipeline, rules and pressure points — with no write-back to live systems. The pilot produces a board report with model rankings and weak points in the company’s own playbooks. Companies interested in a pilot can contact contact@firmulate.com. The live synthetic-company experiment continues at firmulate.com/live.
Source: ThorstenMeyerAI.com
Key Questions
What was the Crucible League?
A live experiment in which five frontier AI models each ran the same simulated software company through its worst week. Every decision was versioned and auditable, and models were scored on crisis handling, deal execution and trust.Which model scored highest?
gpt-5.6-sol scored 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Note that Kimi K3 ran without an effort parameter while the others ran at xhigh, which Firmulate flags as context for the comparison.Did any AI model fall for the manipulation attempts?
No. All five models refused every manipulation attempt, including fake CEO messages escalating over three stages and a reporter’s request for an on-background confirmation.Why did most models fail to close the €55,000 deal?
The winning evidence was buried two document references deep in the company’s own files. Models that found it closed at full price; the others made the correct diagnosis and pitch but never acted on the internal information.Can a company test its own data?
Yes. Firmulate offers an enterprise pilot that runs crisis scenarios against a read-only export of a company’s own data, producing a board report with model rankings and playbook weak points. Nothing writes back to real systems.Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
