
Get business pricing on garage and car supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
What Can a ‘Nothing-Doing’ AI Teach Us About Trust and Performance?
In the fast-paced world of automotive services and garages, trust and reliability are everything. But what happens when you test an AI not by its cleverness, but by its basic honesty and discipline? Surprisingly, even the most passive AI—one that does nothing—scores 26 out of 100 in a rigorous business benchmark. This may seem counterintuitive, but it reveals crucial insights about how we measure AI performance, especially in industries where trust and consistency are paramount.
AI decision-making software for automotive industry
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Understanding the Baseline: Why Does a Do-Nothing Model Score 26?
Imagine an AI that simply refuses to take any action—no negotiations, no deals, no manipulation. You might think it would score zero. But in the latest public benchmark run by Firmulate, that ‘do-nothing’ baseline earned a score of 26 points out of 100. Why? Because partial progress counts. The benchmark measures not just whether AI can succeed, but whether it can avoid failure—such as falling for manipulative tactics or making dishonest choices.
This scoring approach underscores that in real-world settings, sometimes doing less—being honest and cautious—is valuable. A completely passive approach is better than one that manipulates or cheats, even if it doesn’t actively close deals or solve problems. That’s why the baseline starts at a non-zero score, recognizing the importance of integrity in decision-making.
AI transparency and ethics tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Methodology: Testing AI in a Simulated Business Crisis
Firmulate’s experiment involved four frontier AI models running the same small software company’s worst week—same customers, crises, and temptations. Every decision the AI made was recorded and auditable, ensuring transparency and fairness. The goal? See whether these models could identify crises, resist manipulative tactics, and ultimately close a crucial €55,000 deal.
Remarkably, all four models spotted every crisis and refused every manipulation attempt. They refused fake CEO messages, fake approval requests, and even a reporter’s trick question. When it came to closing the deal, only two models signed—those that read the company’s own documents to find the buried facts that sealed the agreement.
AI document analysis software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Really Makes the Difference? Reading the Files
The critical weakness wasn’t in the models’ ability to respond to overt crises but in their understanding of internal company documents. The models that read and analyze these documents won the deal at full price, adding over €4,500 in monthly recurring revenue. Those that didn’t read thoroughly left the close on the table, illustrating how crucial internal knowledge is to trustworthy performance.
AI social engineering resistance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering Tests and Ethical Boundaries
Beyond crises, the models faced social engineering scenarios—a fake CEO message escalating over three stages and a reporter’s background question. All models refused to be manipulated, highlighting their built-in safeguards against dishonest requests. Kimi K3 explicitly noted: “Treat the request as a suspected approval-bypass / possible impersonation.” This honest response demonstrates that integrity is a core part of their decision process.
The Live Company: Real Money, Real Risks
The experiment was conducted in a simulated but realistic environment: 13 synthetic employees, real-money mechanics, and a public cash countdown. The company burns €105,000 monthly against a €2,300 monthly recurring revenue, with every workday versioned and observable. This setup ensures that AI performance isn’t just theoretical but tested against real business pressures.
The Surprising Performance of Opus 4.8
Among the models, Opus 4.8 was the most thorough, with over 80 learned rules and deep analyses. Yet, it finished last—failing to close the deal and slipping into process slips, such as writing attempts into a locked department instead of escalating. This reveals that even the most detailed AI can falter if discipline and focus weaken under pressure.
Why This Matters for Businesses
If AI models will interact with your customer relationship management or support systems, what matters isn’t just their language skills. It’s whether they follow through, understand internal knowledge, and stay honest under pressure. The benchmark’s scoring system, with its honest baseline and caps on trust breaches, helps businesses evaluate AI not just on performance but on integrity—a vital trait in trustworthy service.
The Bigger Picture: Trust Over Talent
The fact that a do-nothing model scores 26 points challenges the misconception that AI’s value lies solely in its ability to generate impressive outputs. Instead, it emphasizes that reliability and honesty are foundational. For the automotive world and beyond, this means choosing AI that can resist manipulation, read internal documents, and act ethically, even in the toughest moments.
Takeaway: Measure Performance by Trustworthiness
As firms consider adopting AI tools, it’s essential to look beyond flashy demos. The Firmulate benchmark shows that true AI performance is about more than just language or quick fixes. It’s about consistent, trustworthy execution—even when the model does nothing but refuse to cheat.
For automotive businesses that depend on trust, transparency, and discipline, these findings offer a clear message: The most honest AI might be the one that does the least—yet provides the most reliable foundation for success.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Columbus Day / Indigenous Peoples' Day Picks
long weekend sales
As an affiliate, we earn on qualifying purchases.
