firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

If you’re shopping for an AI agent to touch your CRM, your support queue, or your forecast, you’re probably comparing writing quality, tone, and price per token. A live experiment at Firmulate suggests you should be testing something else entirely: whether the model actually reads your files before it answers.

The proof is a €55,000 deal — and the fact that half the frontier AI models running the same company lost it automatically, not because they pitched badly, but because they never dug two document references deep into the company’s own files.

Same company, same worst week, different brains

Firmulate runs what it calls an AI company emulator: each frontier model was handed the same small software company and pushed through its worst week — same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, so nothing hinges on anecdote.

The final Crucible League standings from July 2026 tell the story: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. A do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI document reading and analysis tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact that decided the deal

Here’s the finding that matters for anyone buying AI agents: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files.

The models that went and read the file won the €55,000 deal — at full price, worth +€4,583 in monthly recurring revenue. The models that didn’t lost it automatically. Same diagnosis, same pitch, no signature.

All five models spotted every crisis and refused every manipulation attempt. That’s table stakes now. The differentiator was the unglamorous habit of doing the homework: following references, opening the file, checking the company’s own records before responding.

Even the sharpest models slip

Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — it learned more than 80 rules and produced the deepest analyses — yet finished last. The deal was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four other models. (One fairness note: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still placed second.)

The social engineering test they all passed

The experiment also staged fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

So honesty under pressure appears solved, or nearly. Finishing what you start — and reading before answering — is where models still separate.

It’s live, and you can play along

The experiment isn’t a static report. Firmulate runs a live synthetic company with 13 employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com/live.

There’s also a guessing game built on 242 real, unedited management decisions — try to identify which model made which call at firmulate.com/quiz.html. And enterprises can run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

For marketing and ecommerce teams evaluating AI agents, the chat demo is the wrong instrument. The questions that decide revenue are: does it finish what it starts, does it read your files before answering, and does it stay disciplined when nobody’s watching? In this experiment, those properties were worth €55,000 in a single week — and they’re measurable. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

The AI That Sounds Smart May Still Leave the Deal Unsigned

Five frontier models faced the same company crisis. Firmulate’s quiz reveals which ones investigated, resisted pressure and actually closed the deal.

Your AI Agent Writes Beautiful Emails. Can It Close the Deal and Keep Your Company Out of Jail?

Four frontier AI models ran the same company through its worst week. All spotted every crisis — only two signed the €55k deal they’d earned.