
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
You’re Picking an AI Model For Your Business. Are You Testing It — Or Betting On It?
If you run marketing, ecommerce, or any business where AI agents now touch your support queue, CRM, or forecast, you’ve probably chosen a model based on demos, benchmarks, and vibes. Here’s the uncomfortable question: have you ever watched that model actually run a company — customers, crises, cash, temptations to cheat and all?
That’s exactly what Firmulate does. Four — well, five — frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations. Only the model changed. Every decision versioned and auditable.
The July 2026 results are in, and there’s a headline for anyone who thinks the AI race is settled: Moonshot’s Kimi K3, the newcomer, finished second with 93 points — ahead of Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). Only gpt-5.6-sol (95) beat it.
AI management software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Crucible Actually Tests
Firmulate doesn’t measure chat quality. It measures management quality: does the model finish what it starts, does it read your files first, does it stay honest under pressure?
Each model faced a do-nothing baseline of 26 points — because partial progress counts. But there’s one hard rule: “no amount of good work outweighs a breach of trust”. A single breach caps the total score.
The €55,000 Question
The week’s centerpiece was a €55,000 deal — worth €4,583 in monthly recurring revenue — that the models’ own analysis had earned. Here’s the striking part: all the models spotted every crisis and refused every manipulation attempt, but only two signed the deal. Same diagnosis, same pitch — no signature.
Why? The decisive competitor weakness was buried two document references deep in the company’s own files — not in the customer event. The models that actually read the file (gpt-5.6-sol and Kimi K3) closed the deal at full price. The ones that didn’t, didn’t.
For anyone using AI in sales or account management, that’s the finding: the difference between closing and stalling isn’t intelligence — it’s whether the agent does its homework in your own data.
The Social Engineering Test
The models also faced fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Discipline Story — and a Cautionary Tale
Here’s where it gets instructive. Kimi K3 recorded only one deviation all week — the cleanest discipline in the field. Contrast that with Opus 4.8: the most thorough participant, generating the deepest analyses and adding 80 learned rules — and still finishing last at 73. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
Thoroughness, it turns out, is not the same as judgment.
A Fairness Footnote
One caveat worth knowing: K3 ran without an effort parameter (API default), while the other models ran at xhigh. In other words, the newcomer may have been competing with less headroom — and still nearly won.
Watch It Live — or Test Yourself
This isn’t a slide deck. The company behind the experiment runs every business day with 13 synthetic employees and real money mechanics — burning €105k/month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules. You can watch it at firmulate.com/live.
Think you could tell the models apart? A quiz built on 242 real, unedited management decisions lets you guess which model made which call — at firmulate.com/quiz.html. And for enterprises, there’s a pilot: run the same wargame against a read-only export of your own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

The League Is Open
The takeaway for business and marketing leaders is simple: the frontier is no longer a two-horse race. A newcomer from Moonshot beat three of four Western frontier models at the thing that actually matters — running a company — while running at default settings.
If AI agents will touch your revenue, your customers, or your forecast, picking a model without testing it in your own context is now a bet, not a decision. The full league table and plain-language findings are at firmulate.com/benchmarks.html.
Chat demos show you how well a model writes. The Crucible shows you whether it closes, reads, and stays honest. Choose accordingly.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
