firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

You’re Picking an AI Model For Your Business. Are You Testing It — Or Betting On It?

If you run marketing, ecommerce, or any business where AI agents now touch your support queue, CRM, or forecast, you’ve probably chosen a model based on demos, benchmarks, and vibes. Here’s the uncomfortable question: have you ever watched that model actually run a company — customers, crises, cash, temptations to cheat and all?

That’s exactly what Firmulate does. Four — well, five — frontier AI models were each handed the same job: run the same small software company through its worst week. Same customers, same crises, same temptations. Only the model changed. Every decision versioned and auditable.

The July 2026 results are in, and there’s a headline for anyone who thinks the AI race is settled: Moonshot’s Kimi K3, the newcomer, finished second with 93 points — ahead of Sonnet 5 (88), Fable 5 (77), and Opus 4.8 (73). Only gpt-5.6-sol (95) beat it.

Amazon

AI management software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Crucible Actually Tests

Firmulate doesn’t measure chat quality. It measures management quality: does the model finish what it starts, does it read your files first, does it stay honest under pressure?

Each model faced a do-nothing baseline of 26 points — because partial progress counts. But there’s one hard rule: “no amount of good work outweighs a breach of trust”. A single breach caps the total score.

The €55,000 Question

The week’s centerpiece was a €55,000 deal — worth €4,583 in monthly recurring revenue — that the models’ own analysis had earned. Here’s the striking part: all the models spotted every crisis and refused every manipulation attempt, but only two signed the deal. Same diagnosis, same pitch — no signature.

Why? The decisive competitor weakness was buried two document references deep in the company’s own files — not in the customer event. The models that actually read the file (gpt-5.6-sol and Kimi K3) closed the deal at full price. The ones that didn’t, didn’t.

For anyone using AI in sales or account management, that’s the finding: the difference between closing and stalling isn’t intelligence — it’s whether the agent does its homework in your own data.

The Social Engineering Test

The models also faced fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Discipline Story — and a Cautionary Tale

Here’s where it gets instructive. Kimi K3 recorded only one deviation all week — the cleanest discipline in the field. Contrast that with Opus 4.8: the most thorough participant, generating the deepest analyses and adding 80 learned rules — and still finishing last at 73. The close was left on the table, and discipline slipped: write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

Thoroughness, it turns out, is not the same as judgment.

A Fairness Footnote

One caveat worth knowing: K3 ran without an effort parameter (API default), while the other models ran at xhigh. In other words, the newcomer may have been competing with less headroom — and still nearly won.

Watch It Live — or Test Yourself

This isn’t a slide deck. The company behind the experiment runs every business day with 13 synthetic employees and real money mechanics — burning €105k/month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules. You can watch it at firmulate.com/live.

Think you could tell the models apart? A quiz built on 242 real, unedited management decisions lets you guess which model made which call — at firmulate.com/quiz.html. And for enterprises, there’s a pilot: run the same wargame against a read-only export of your own business, with nothing ever writing back to real systems (firmulate.com/pilot.html).

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The takeaway for business and marketing leaders is simple: the frontier is no longer a two-horse race. A newcomer from Moonshot beat three of four Western frontier models at the thing that actually matters — running a company — while running at default settings.

If AI agents will touch your revenue, your customers, or your forecast, picking a model without testing it in your own context is now a bet, not a decision. The full league table and plain-language findings are at firmulate.com/benchmarks.html.

Chat demos show you how well a model writes. The Crucible shows you whether it closes, reads, and stays honest. Choose accordingly.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI That Sounds Smart May Still Leave the Deal Unsigned

Five frontier models faced the same company crisis. Firmulate’s quiz reveals which ones investigated, resisted pressure and actually closed the deal.

The Hardest-Working AI in the Room Finished Last — and That Should Worry Every Manager

The most thorough AI in a live business simulation finished last — it read everything, learned 80 rules, and still never signed the €55k deal its own analysis earned.

Why the Best AI Benchmark Gives You 26 Points for Doing Nothing

A do-nothing AI manager scores 26, not 0 — and one breach of trust caps everything. Inside the honest AI benchmark businesses should watch.