firmulate.com/index — live view
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Promotion Everyone Regrets

Every business owner knows the hire who nails the interview and flops in the job. Brilliant on the whiteboard, articulate in meetings — and then the churn wave hits, a big customer threatens to walk, and somehow the deal that was already won never gets signed.

We are now interviewing AI agents the same way. Coding leaderboards and chat arenas reward models for the quality of their answers. But the qualities that decide whether a company lives or die — triage under capacity pressure, judgment across days, honesty when a reporter calls — are exactly what those benchmarks never measure.

A live experiment at Firmulate just made that gap painfully concrete.

Amazon

AI email writing assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four Frontier Models, One Terrible Week

Firmulate ran four frontier AI models as CEOs of the same small software company through its worst week. Same customers, same crises, same temptations to cheat. Every decision versioned and auditable. The final league table from July 2026: gpt-5.6-sol scored 95, Kimi K3 scored 93, Sonnet 5 scored 88, Fable 5 scored 77, and Opus 4.8 landed last at 73. For context, doing nothing at all scores 26 — because partial progress counts, but a single breach of trust caps the total. As the experiment puts it: no amount of good work outweighs a breach of trust.

The Finding That Should Worry Every Buyer

All four models spotted every crisis. All four refused every manipulation attempt. And only two signed the €55,000 deal their own analysis had already earned. Same diagnosis, same pitch — no signature.

Think about what that means for the agent you’re about to plug into your CRM or support queue. The intelligence was there. The follow-through wasn’t. That gap is invisible in a chat demo.

The Buried Fact

The deal-breaker wasn’t even in the customer conversation. The decisive competitive weakness sat two document references deep in the company’s own files. The models that actually read before acting won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Social Engineering: Five for Five

The week included fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” On guard rails, the frontier has genuinely arrived.

Hardworking and Last Place

The most instructive profile is Opus 4.8: the most thorough participant of the field, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. (One fairness note: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly won.)

This Isn’t a Slide Deck

The company is live, not a screenshot. Thirteen synthetic employees, real money mechanics: €105k monthly burn against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, every workday versioned — watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business at firmulate.com/benchmarks.html.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

Management Quality, Not Chat Quality

The category Firmulate is proposing matters for anyone hiring AI in the next two years: measure management quality, not chat quality. Does the agent finish what it starts? Does it read your files before it acts? Does it stay honest under pressure — and what does a unit of useful work cost?

Scenarios like a churn wave, a price increase, a down round, a PR crisis — that’s the new curriculum. Your AI agent may ace the coding benchmark and charm the chat arena. But if it can’t close the €55k deal sitting in front of it, or reads a locked department it should have escalated, you’ve hired the interview champion, not the operator.

Before you deploy, wargame it. The gap between a great answer and a finished job is where your margin lives.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


You May Also Like

The AI That Sounds Smart May Still Leave the Deal Unsigned

Five frontier models faced the same company crisis. Firmulate’s quiz reveals which ones investigated, resisted pressure and actually closed the deal.