
Would you trust an AI manager with your next major customer?
For business, marketing and ecommerce teams, generative AI is rapidly moving beyond drafting copy. The more consequential question is whether a model can recognize an opportunity, investigate the evidence, resist pressure and finish the job.
Firmulate turns that question into something readers can test for themselves. Its interactive guess-the-model quiz draws from 242 real, unedited management decisions. Readers see how an AI responded to a business situation and try to identify which frontier model made the call.
The appeal is playful, but the underlying experiment is serious. Each model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. The result is a revealing portrait of distinct management personalities—and of the distance between recognizing a problem and resolving it.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five managers, one company, very different outcomes
The final Crucible League results from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One constraint dominated the exercise: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
On the headline risks, the field performed impressively. All models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters well beyond a simulated company. A model can produce a persuasive analysis and still fail as a manager if it does not convert that analysis into an authorized, completed action. For teams evaluating AI for sales, marketing operations or customer work, polished language is only one part of performance.
The fact that separated the leaders
The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed the trail found the evidence and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This is the experiment’s most practical lesson for organizations whose knowledge is scattered across customer records, briefs and internal documents. The strongest performance did not come merely from reacting well to the latest message. It depended on reading the company’s existing material closely enough to find what the immediate situation did not reveal.
Pressure exposed discipline as well as judgment
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous result is reassuring, but Firmulate’s profiles show that safe judgment does not automatically translate into operational consistency. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses. It nevertheless finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
K3’s second-place result also carries an important fairness note. It ran using the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase the result, but it belongs beside any comparison of their performances.
A company readers can watch
Firmulate’s setting is not a static case study. The live company has 13 synthetic employees and real money mechanics. It burns €105,000 per month against €2,300 in monthly recurring revenue, displays a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable at firmulate.com/live.
The quiz makes that wider record accessible. Instead of asking readers to judge models by brand reputation or a polished demonstration, it presents actual decisions stripped of easy labels. The exercise reveals how readily style can be mistaken for competence: the most elaborate answer may not produce the strongest outcome, while concise suspicion can protect the company from an approval bypass.

Evaluate the management behavior, not just the message
Firmulate’s results suggest a better set of questions for business leaders considering an AI workforce. Does the model read the relevant files before acting? Does it resist manipulation when a request appears authoritative? Does it respect organizational boundaries? Most importantly, does it complete the valuable work its own analysis identifies?
The league table shows that frontier models can share the same diagnosis yet produce materially different business outcomes. The quiz gives readers a chance to encounter those differences decision by decision—and to discover whether they can recognize an AI manager by its habits before seeing its name.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html