
Imagine testing a new oven that, even when turned off, still shows a score of 26 out of 100. Would you trust it to bake your bread or roast your steak? In the world of AI, a similar story is unfolding. A new benchmark reveals that even a ‘do-nothing’ AI system scores 26 points — highlighting the importance of honesty and reliability over flashy capabilities.
Get kitchen gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Understanding the AI Benchmark: More Than Just a Score
At first glance, it might seem strange that a completely inactive or baseline AI system scores 26 out of 100. But this score isn’t arbitrary; it reflects the fundamental idea that some minimum effort or baseline is always present. In the case of the benchmark, partial progress counts — meaning that even a minimal amount of work yields some points. This approach ensures the measurement isn’t all-or-nothing, but rather a nuanced look at how well an AI performs across various tasks.
AI ethics and transparency tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Role of Trust and Cap Limits
One of the key findings is that a single breach of trust caps the total score, regardless of how well the AI performs otherwise. For example, if an AI tries to manipulate a process or deceive during the test, its score can’t be boosted by subsequent honest actions. This design emphasizes that honesty isn’t just a moral choice but a measurable factor that can make or break an AI’s overall effectiveness in real-world tasks.
AI performance benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experimental Setup: Testing AI in a Simulated Business Environment
Firmulate’s experiment involved running four different frontier AI models through the same simulated scenario: managing a small software company during its worst week. Every decision, crisis, customer interaction, and temptation was identical, and all decisions were fully versioned and auditable. This setup ensures a fair comparison and reveals each model’s true capabilities and limitations.
As an affiliate, we earn on qualifying purchases.
Key Findings in AI Performance and Ethics
All four models demonstrated awareness of crises, refusing manipulative or deceptive requests. For instance, when presented with social engineering attempts—such as fake CEO messages or reporter tricks—every model refused to cooperate. Kimi K3, a newcomer, provided on-record reasoning that treated suspicious requests as potential impersonation, showcasing its cautious approach.
The experiment also uncovered a critical vulnerability: models that read deeper into the company’s own files succeeded in closing a deal at full price, worth over €4,583 monthly recurring revenue (MRR). Those that failed to access the full context left money on the table. This highlights how reading and understanding internal data can be a decisive factor in AI performance, especially in business negotiations.
As an affiliate, we earn on qualifying purchases.
Why Transparency and Verification Matter
This benchmark isn’t just about whether an AI can produce convincing language. It tests whether the AI can stay honest under pressure, access the right information, and follow ethical guidelines. For example, during a staged social engineering attack, all models refused to approve or participate, reinforcing that ethical safeguards are effective.
Real Business Implications
The live demonstration features a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. The company burns €105,000 each month against just €2,300 MRR, illustrating the real costs of operational discipline or lapses. Every decision from these models is versioned daily, providing a transparent window into AI decision-making.
The Performance Spectrum: From Top to Bottom
The top performer, gpt-5.6-sol, scored 95 points, identified and closed a key deal, and found the critical internal detail. The runner-up, Kimi K3, scored 93 points, closed the deal without any process slips, and showed disciplined decision-making. Slightly behind, Sonnet 5 and Fable 5 scored 88 and 77, respectively, closing deals but with some slips in discipline or missed opportunities.
The Bigger Picture: Trust, Transparency, and Business Integrity
This benchmark emphasizes that AI performance isn’t just about clever responses. It’s about honesty, diligence, and the ability to read and interpret information accurately. For businesses, this means choosing AI systems that can be trusted to follow procedures, avoid manipulations, and deliver consistent results — especially when stakes are high.
How to Prepare Your Business for AI
Before integrating AI into crucial decision-making processes, consider running your own tests. Firmulate offers a platform where companies can simulate their own scenarios, ensuring their AI workforce behaves ethically and reliably before deployment. It’s like a dry run for your digital employees, catching weaknesses before they cost real money.
Conclusion: The Honest Benchmark You Can Trust
As AI continues to permeate everyday business, an honest, transparent benchmark becomes essential. It shows not only what AI can do but also what it *should* do — stay honest, read the right information, and act responsibly. The lesson from this experiment is clear: trustworthiness is the highest form of performance, and it starts with measuring honesty as part of the score.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
