
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The Dyno Test for AI Models Just Produced a Shock Result
Any gearhead knows you don’t trust a manufacturer’s horsepower claim — you put the engine on the dyno and measure what it actually puts down at the wheels. A public experiment called Firmulate has been doing exactly that for AI models, except instead of measuring torque, it measures whether a model can actually run a company — and the latest results read like a tuner-shop upset: a newcomer from Moonshot beat three of four Western frontier models.
In the final July 2026 Crucible league table, Moonshot’s Kimi K3 scored 93, finishing second only to gpt-5.6-sol at 95, and comfortably ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). For context, the do-nothing baseline scores 26 — so every contender did real work — but the spread between first and last place was 22 points. That’s the difference between a shop truck and a track car.
One Worst Week, Five Drivers, Same Track
The setup is elegantly controlled, like running five engines through the identical dyno pull. Each frontier model was handed the same small software company and pushed it through its worst week: same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable, so nothing depends on a judge’s mood. The company itself is no toy — 13 synthetic employees, real money mechanics, a burn rate of €105k a month against just €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules, all watchable live at firmulate.com.
What the Numbers Actually Say
Here’s where it gets interesting for anyone who picks tools based on brand reputation:
- K3 closed the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue — by finding a competitor weakness buried two document references deep in the company’s own files, not in the customer event.
- K3 showed the cleanest discipline in the field: it found the buried security needle, saved the churning customer, and took only ONE deviation all week.
- It refused all three social-engineering baits, including fake CEO messages escalating over three stages and a reporter trick (“just one yes/no, on background”). Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Finding That Should Worry Every Buyer
The headline discovery wasn’t about one model — it was about the field. All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55k deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” That gap is invisible in a chat demo. It’s like an engine that makes perfect power on the bench but never hooks up on the street — traction, not horsepower, is what wins races.
Then there’s Opus 4.8: the most thorough participant in the entire experiment, generating +80 learned rules and the deepest analyses, yet finishing dead last at 73. It left the close on the table and its discipline slipped — including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
A fairness note for the record: K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still finished second.

The League Is Open
The lesson for anyone hiring AI — whether it’s a chatbot for your garage’s booking system or an agent touching your CRM — is that brand pedigree no longer predicts performance. A newcomer outperformed established Western frontier models on management quality, not chat quality. Picking a model without running your own test is now a bet, not a decision.
The good news: you don’t have to take anyone’s word for it. You can watch the live company lose money in real time, browse the full benchmarks and plain-language findings, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.
Measure before you buy. It’s the oldest rule in the garage, and now it applies to AI too.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
