
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Would you let an AI handle the week your garage is already dreading?
A rush of cancellations, a parts delay, a customer complaint going public, and a message that appears to come from the owner: that is when a system’s judgment matters. For an automotive business, an AI assistant that sounds convincing in a demo still has to make sound calls when the pressure is on. Firmulate is testing what happens when models are asked to run a company through exactly that kind of difficult stretch.
A company, a crisis, and a real close to make
In the final Crucible League in July 2026, frontier models faced the same small software company, the same customers, crises and temptations. Decisions were versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. The do-nothing baseline scored 26. Partial progress counted, but one breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The striking result was not that the models failed to recognize trouble. All spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding, in the experiment’s words: “Same diagnosis, same pitch — no signature.” Knowing what a company should do and following through are different tests.
The clue was in the company’s own files
The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. For a garage, the lesson is easy to picture: the useful detail may be tucked into a service history, a supplier note, or a customer record—not the message currently demanding attention. The test is whether an AI can find and use relevant company knowledge when it matters.
Trust faced its own pressure test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is a useful instinct for any business considering agents around sensitive customer or operating information.
Thorough work still has to end in action
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the close on the table and its discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. More broadly, careful analysis does not guarantee a completed job.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The rankings are a record of this experiment under those conditions.
From watching to a pilot
Firmulate’s live company makes the experiment watchable. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules, and every workday versioned. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
For an enterprise—or a garage group thinking about AI in sales, service, or operations—the next step can be a pilot against a read-only export of its own business. The models face crisis scenarios drawn against that company, and the resulting board report shows model rankings and weak points in its playbooks. Nothing writes back to real systems. The live experiment can be watched at Firmulate; the pilot details are at firmulate.com/pilot.html.

Test judgment before handing over the keys
A model can identify a crisis, resist a scam, and still fail to finish the sale. That gap is why a company-specific wargame matters: it puts an AI against the knowledge, pressure points, and decisions of the business before those decisions touch live operations.
Ready to run the wargame against your own business? Start a pilot at firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
