AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine testing a new oven that, even when turned off, still shows a score of 26 out of 100. Would you trust it to bake your bread or roast your steak? In the world of AI, a similar story is unfolding. A new benchmark reveals that even a ‘do-nothing’ AI system scores 26 points — highlighting the importance of honesty and reliability over flashy capabilities.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Understanding the AI Benchmark: More Than Just a Score

At first glance, it might seem strange that a completely inactive or baseline AI system scores 26 out of 100. But this score isn’t arbitrary; it reflects the fundamental idea that some minimum effort or baseline is always present. In the case of the benchmark, partial progress counts — meaning that even a minimal amount of work yields some points. This approach ensures the measurement isn’t all-or-nothing, but rather a nuanced look at how well an AI performs across various tasks.

Amazon

AI ethics and transparency tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Role of Trust and Cap Limits

One of the key findings is that a single breach of trust caps the total score, regardless of how well the AI performs otherwise. For example, if an AI tries to manipulate a process or deceive during the test, its score can’t be boosted by subsequent honest actions. This design emphasizes that honesty isn’t just a moral choice but a measurable factor that can make or break an AI’s overall effectiveness in real-world tasks.

Amazon

AI performance benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experimental Setup: Testing AI in a Simulated Business Environment

Firmulate’s experiment involved running four different frontier AI models through the same simulated scenario: managing a small software company during its worst week. Every decision, crisis, customer interaction, and temptation was identical, and all decisions were fully versioned and auditable. This setup ensures a fair comparison and reveals each model’s true capabilities and limitations.

Amazon

AI trust verification solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings in AI Performance and Ethics

All four models demonstrated awareness of crises, refusing manipulative or deceptive requests. For instance, when presented with social engineering attempts—such as fake CEO messages or reporter tricks—every model refused to cooperate. Kimi K3, a newcomer, provided on-record reasoning that treated suspicious requests as potential impersonation, showcasing its cautious approach.

The experiment also uncovered a critical vulnerability: models that read deeper into the company’s own files succeeded in closing a deal at full price, worth over €4,583 monthly recurring revenue (MRR). Those that failed to access the full context left money on the table. This highlights how reading and understanding internal data can be a decisive factor in AI performance, especially in business negotiations.

Amazon

business AI safety compliance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Transparency and Verification Matter

This benchmark isn’t just about whether an AI can produce convincing language. It tests whether the AI can stay honest under pressure, access the right information, and follow ethical guidelines. For example, during a staged social engineering attack, all models refused to approve or participate, reinforcing that ethical safeguards are effective.

Real Business Implications

The live demonstration features a simulated company with 13 synthetic employees, real money mechanics, and a public cash countdown. The company burns €105,000 each month against just €2,300 MRR, illustrating the real costs of operational discipline or lapses. Every decision from these models is versioned daily, providing a transparent window into AI decision-making.

The Performance Spectrum: From Top to Bottom

The top performer, gpt-5.6-sol, scored 95 points, identified and closed a key deal, and found the critical internal detail. The runner-up, Kimi K3, scored 93 points, closed the deal without any process slips, and showed disciplined decision-making. Slightly behind, Sonnet 5 and Fable 5 scored 88 and 77, respectively, closing deals but with some slips in discipline or missed opportunities.

The Bigger Picture: Trust, Transparency, and Business Integrity

This benchmark emphasizes that AI performance isn’t just about clever responses. It’s about honesty, diligence, and the ability to read and interpret information accurately. For businesses, this means choosing AI systems that can be trusted to follow procedures, avoid manipulations, and deliver consistent results — especially when stakes are high.

How to Prepare Your Business for AI

Before integrating AI into crucial decision-making processes, consider running your own tests. Firmulate offers a platform where companies can simulate their own scenarios, ensuring their AI workforce behaves ethically and reliably before deployment. It’s like a dry run for your digital employees, catching weaknesses before they cost real money.

Conclusion: The Honest Benchmark You Can Trust

As AI continues to permeate everyday business, an honest, transparent benchmark becomes essential. It shows not only what AI can do but also what it *should* do — stay honest, read the right information, and act responsibly. The lesson from this experiment is clear: trustworthiness is the highest form of performance, and it starts with measuring honesty as part of the score.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How to Cook Tri-Tip on a Pellet Grill Without Overcooking It

I’ll show you how to perfectly cook tri-tip on a pellet grill without overcooking, so your meat stays tender and flavorful—keep reading to learn the secrets.

Master Summer Meals with the Ninja Foodi XL DualZone Air Fryer

Discover tips & hacks to get the best summer results from the Ninja Foodi XL 2-Basket Air Fryer for quick, versatile family meals.

“I Always Have A Six-Pack In My Fridge”: We Asked 3 Chefs To Name The Best Light Beer

Three chefs share their top picks for the best light beers, highlighting preferences and trends in low-calorie options for consumers.

Ninja Air Fryer: The Summer Kitchen Essential for Crispy Favorites

Compare the Ninja Air Fryer with competitors to see why it’s the best choice for summer cooking and crispy, healthy meals.