
Imagine you’re running a busy kitchen, and your trusted oven suddenly malfunctions during a dinner rush. You need a quick fix, but more importantly, you need to trust that your tools will perform under pressure—not just in perfect demo conditions, but during the actual chaos. The same principle applies to AI in business. While many AI chatbots impress in isolated demos, can they handle the messy, high-stakes moments that truly test management quality? That’s what a groundbreaking live experiment from Firmulate sets out to discover.
What is Firmulate’s Live Business Wargame?
Firmulate runs a real-time, fully transparent simulation where AI models are tasked with managing a small software company facing its worst week. This isn’t a staged demo or a simple Q&A—it’s a complex scenario involving real crises, customer demands, internal temptations, and strategic decisions. Each model operates a synthetic but realistic business environment, complete with 680+ self-learned rules, a cash countdown, and actual revenue mechanics.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Testing Ground: Real Crises, Real Decisions
All four frontier AI models—ranging from the well-known GPT-5.6-sol to newer entrants like Kimi K3—were given the same challenges. They had to navigate customer crises, read critical internal files, and refuse manipulative requests like fake CEO messages. Every decision was auditable, every move recorded, reflecting the true management quality of each AI.
Key Findings: Capable of Spotting Crises but Not Always Closing Deals
Remarkably, all four models identified every crisis and refused every manipulation attempt. They proved effective at recognizing threats and maintaining honesty under pressure. However, when it came to closing business deals, only two signed the €55,000 deal their own analysis had earned. The others, despite diagnosing correctly, left the opportunity on the table due to process slips or discipline lapses.
business crisis simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Hidden Weaknesses in Deep Files
The decisive factor in winning a deal was reading and understanding internal documents—two document references deep in the company’s files. Models that accessed and interpreted these files secured the full deal, worth an additional €4,583 in monthly recurring revenue (MRR). This shows that true management quality isn’t just about surface-level answers but understanding the full context and nuances of internal information.
AI ethical decision support systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Social Engineering and Ethical Vigilance
In a test of ethical resilience, all models refused to engage with fake CEO messages staged over multiple stages plus a reporter trick. Kimi K3 explicitly justified its refusal by treating the request as a suspected impersonation or approval bypass. This demonstrates that current AI, when properly designed, can uphold ethical boundaries even under simulated social engineering pressure.
internal document analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Human-Like Business: Real Money, Real Losses
Underlying this experiment is a live, functioning company with 13 synthetic employees managing real money mechanics—burning €105k monthly against €2.3k MRR. The company operates every business day, with every decision versioned for transparency. The live site, firmulate.com/live, offers an ongoing view into this ongoing test, making it a rare window into AI’s practical capabilities and limitations in complex, real-world scenarios.
Why Management Quality Matters More Than Chat Skills
This experiment highlights a crucial insight: the traditional benchmarks and chat-focused demos don’t reveal whether AI can deliver management quality under stress. It’s one thing to generate convincing responses; it’s another to read internal files, maintain discipline, recognize manipulative social engineering, and close deals—especially in a high-pressure environment with real money at stake.
What This Means for Business Leaders
Ultimately, the question isn’t whether AI can pass isolated chat tests or answer questions correctly. It’s whether AI can finish what it starts, stay honest, understand complex internal context, and deliver tangible results—especially during crises. As firms consider deploying AI into CRM, customer support, or forecasting, these are the metrics that will matter most.
The League Table of Management Performance
- gpt-5.6-sol 95: Identified buried facts, closed the deal, demonstrating full management capability.
- Kimi K3 93: Clean discipline, signed the deal, but ran without an effort parameter.
- Sonnet 88: Signed the deal but showed some process slips.
- Sonnet 77: Also signed but with more discipline slips.
These results, from the live experiment, set a new bar for understanding what AI can do in real-world management—not just chat quality but actual decision-making under pressure.
Try It Yourself
Businesses interested in testing their own AI tools can run similar simulations against their data, without risking real systems. Visit firmulate.com/pilot.html to see how a wargame can help you assess your AI’s readiness before deployment.

Real-world AI management isn’t about chat quality; it’s about decision-making under pressure, understanding internal context, and staying honest. Firmulate’s live experiment proves that the true measure of AI readiness requires testing in high-stakes, realistic scenarios—so your AI can perform when it matters most.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html