AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on garage and car supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The Dyno Test for AI Models Just Produced a Shock Result

Any gearhead knows you don’t trust a manufacturer’s horsepower claim — you put the engine on the dyno and measure what it actually puts down at the wheels. A public experiment called Firmulate has been doing exactly that for AI models, except instead of measuring torque, it measures whether a model can actually run a company — and the latest results read like a tuner-shop upset: a newcomer from Moonshot beat three of four Western frontier models.

In the final July 2026 Crucible league table, Moonshot’s Kimi K3 scored 93, finishing second only to gpt-5.6-sol at 95, and comfortably ahead of Sonnet 5 (88), Fable 5 (77) and Opus 4.8 (73). For context, the do-nothing baseline scores 26 — so every contender did real work — but the spread between first and last place was 22 points. That’s the difference between a shop truck and a track car.

One Worst Week, Five Drivers, Same Track

The setup is elegantly controlled, like running five engines through the identical dyno pull. Each frontier model was handed the same small software company and pushed it through its worst week: same customers, same crises, same temptations to cut corners. Every decision is versioned and auditable, so nothing depends on a judge’s mood. The company itself is no toy — 13 synthetic employees, real money mechanics, a burn rate of €105k a month against just €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules, all watchable live at firmulate.com.

What the Numbers Actually Say

Here’s where it gets interesting for anyone who picks tools based on brand reputation:

  • K3 closed the €55,000 deal at full price — worth +€4,583 in monthly recurring revenue — by finding a competitor weakness buried two document references deep in the company’s own files, not in the customer event.
  • K3 showed the cleanest discipline in the field: it found the buried security needle, saved the churning customer, and took only ONE deviation all week.
  • It refused all three social-engineering baits, including fake CEO messages escalating over three stages and a reporter trick (“just one yes/no, on background”). Its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Finding That Should Worry Every Buyer

The headline discovery wasn’t about one model — it was about the field. All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55k deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” That gap is invisible in a chat demo. It’s like an engine that makes perfect power on the bench but never hooks up on the street — traction, not horsepower, is what wins races.

Then there’s Opus 4.8: the most thorough participant in the entire experiment, generating +80 learned rules and the deepest analyses, yet finishing dead last at 73. It left the close on the table and its discipline slipped — including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

A fairness note for the record: K3 ran without an effort parameter (API default) while the other models ran at xhigh — and still finished second.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

The lesson for anyone hiring AI — whether it’s a chatbot for your garage’s booking system or an agent touching your CRM — is that brand pedigree no longer predicts performance. A newcomer outperformed established Western frontier models on management quality, not chat quality. Picking a model without running your own test is now a bet, not a decision.

The good news: you don’t have to take anyone’s word for it. You can watch the live company lose money in real time, browse the full benchmarks and plain-language findings, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Measure before you buy. It’s the oldest rule in the garage, and now it applies to AI too.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Launch Control Explained

Boost your understanding of launch control and discover how it optimizes starts—keep reading to uncover the secrets behind this high-performance technology.

A Serious Injury Kept Me Off Motorcycles. Honda’s E-Clutch Got Me Back in the Saddle

A rider recovering from a serious injury credits Honda’s E-Clutch technology for enabling their return to motorcycling, highlighting advancements in safety and accessibility.

Traffic Regulators Vow to Unleash ‘Innovation’ by Eliminating Robotaxi Brake Pedals

Regulators aim to eliminate brake pedals in robotaxis to foster innovation, citing safety and technological advancements. Details on implementation remain unclear.

Mazda, Once the Loudest Critic of Touchscreens, Now Says They’re Safer Than Buttons: TDS

Mazda, previously critical of touchscreens, now states they are safer than traditional buttons, citing safety benefits and technological advancements.