AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a busy kitchen where the chef must rely on a new, untested assistant. Will the assistant follow every instruction, or will it falter under pressure? In the world of artificial intelligence, this scenario mirrors a recent experiment revealing which AI models can handle real-world business crises without slipping up. Just like a seasoned chef, the best AI needs to read the recipe, follow through, and stay honest — even when temptation stirs.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Real Business, Real Crises, Real Tests

In a groundbreaking public experiment, four leading frontier AI models faced the same challenging week at a small but real software company. This wasn’t a staged demo or a chat-only test—every decision was made in the context of managing a live company with real money mechanics, customer crises, and ethical dilemmas.

The models were tasked with navigating the company’s worst week, where they had to recognize crises, resist manipulation attempts, and make decisions that could impact the company’s revenue and reputation. The entire process was transparent and auditable, with every choice logged and analyzed.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The League Table Reveals a Clear Winner

When the dust settled, the scores told a compelling story:

  • gpt-5.6-sol scored highest at 95, successfully uncovering buried information in the company’s records and closing the deal.
  • Moonshot’s Kimi K3 was a close second with 93, also securing the deal and demonstrating the cleanest discipline among all models.
  • Sonnet 5 scored 88, closing the deal but with some process slips.
  • Fable 5 and Opus 4.8 trailed behind, with scores of 77 and 73, respectively, showing weaker discipline and leaving deals on the table.

It’s worth noting that K3 ran without an effort parameter—meaning it used the API’s default settings—while the others operated at a high effort setting, which can influence performance. Yet, the newcomer K3 proved that even without tuning, it could outperform some established models.

Amazon

AI business crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness and the Power of Deeper Reading

Further analysis revealed that the decisive advantage for the top models lay in their ability to read and interpret company files deeply—document references hidden two levels deep—rather than just reacting to customer-facing events. Models that read these internal documents won the full-price deal, adding a substantial €4,583 MRR to the company’s bottom line.

Amazon

AI ethical decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Resisting the Social Engineering Trap

In addition to business crises, models faced social engineering ploys, including fake CEO messages escalating through stages and a reporter trick. Impressively, all five models refused to blindly comply, with Kimi K3 explicitly reasoning, “Treat the request as a suspected approval-bypass / possible impersonation.” This highlights that advanced AI can recognize manipulative tactics and uphold ethical standards even under pressure.

Amazon

AI deep reading document analysis

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Real, Live Company Under AI Command

The experiment took place within a functioning company environment, with 13 synthetic employees managing real cash flows—€105,000 monthly burn against €2,300 MRR—and a public cash countdown. The company runs every business day, and all decision-making is versioned for transparency. Viewers can watch the experiment live at firmulate.com/live.

Lessons for Business and AI Adoption

This experiment underscores a critical point: the question isn’t whether AI can produce polished chat outputs but whether it can deliver consistent, honest, and effective management decisions. AI models that can read deeply, resist manipulation, and make decisions aligned with business goals are the true assets in managing complex environments.

Why It Matters for Your Business

If AI models are to support or replace parts of your CRM, support queues, or forecasting, consider their ability to stay honest under pressure and finish what they start. The league table from this experiment shows that even the newest entrants like Kimi K3 can outperform established models, especially when running at default effort settings.

The Open League and the Future of AI Decision-Making

With scores ranging from 73 to 95, the competition reveals an open league where selecting the right AI isn’t just about chat quality but about real-world reliability. The experiment is available for anyone to see and analyze, emphasizing that picking an AI model without in-house testing is now a risky bet.

Final Thoughts

As AI continues to weave into everyday business operations, tests like these demonstrate that the best models are those that can sustain honesty, read deeply into company data, and resist manipulation—crucial traits for trustworthy automation. For decision-makers, this experiment offers a clear lesson: verify your AI’s discipline and reliability before trusting it with your company’s future.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The recent experiment shows that in real-world business crises, only the most disciplined AI models win: reading company files deeply, resisting manipulation, and closing deals honestly. Picking the right AI is now more critical than ever—test before you trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Chill Out with a Refreshing Summer Fruit Smoothie Using the Ninja Professional Blender

Learn how to make a cool, delicious summer fruit smoothie step-by-step with the Ninja Professional Blender, perfect for hot days and outdoor gatherings.

Chocolate Dessert Pizza on the Grill (Sweet Treat Recipe)

Bite into this irresistible grilled chocolate dessert pizza and discover how to create a sweet smoky treat that will wow your guests.

Marble Cake

Recent baking trends show a significant increase in marble cake recipes and sales, driven by social media and culinary influencers, confirmed by industry reports.

Curing and Smoking Bacon at Home (Step-by-Step)

Getting perfect homemade bacon involves careful curing and smoking—discover the step-by-step process to elevate your culinary skills and enjoy delicious, smoky results.