
Imagine trying to run a business during its worst week—crises piling up, temptations to cut corners, and every decision scrutinized under pressure. Now, what if your AI assistant had to do the same? Would it be honest? Would it follow through? Welcome to a groundbreaking live experiment that tests the management personalities of leading AI models in real-world business scenarios.
The Live AI Business Simulator
Firmulate’s experiment is unlike any other. It places four frontier AI models—each with distinct capabilities—inside a real, functioning software company. This isn’t a simulation in the abstract; it’s the company’s actual operations on a typical workday, but with one twist: every decision is recorded, auditable, and made by AI models facing identical crises, temptations, and incentives.
Same Crises, Different Decisions
Each AI model was tasked with navigating the company’s toughest week. Customers demanded urgent fixes, internal conflicts emerged, and there’s always the temptation to bypass procedures for quick results. The models had to identify issues, read critical files, and decide whether to sign deals or escalate issues. Despite identical scenarios, their decision-making personalities revealed striking differences.
Results Speak Louder Than Words
All four models detected every crisis and refused manipulation attempts, such as fake CEO messages or reporter tricks designed to test their integrity. But, when it came to sealing a crucial deal worth €55,000, only two of them signed—the others, despite their accurate analysis, left full value on the table.
Specifically, the ‘gpt-5.6-sol’ model scored the highest at 95 points for its comprehensive performance, including uncovering a hidden document in the company’s files that clinched the deal. The ‘Kimi K3’ model scored 93, also closing the deal, and was lauded for its discipline. Meanwhile, ‘Sonnet 5’ and ‘Fable 5’ scored 88 and 77 respectively; both signed deals but with slips in process discipline.
What Made the Difference?
The key to winning the full deal lay in reading deeper into company documents. The models that examined internal files—beyond just the customer interactions—uncovered critical information that justified full price negotiations. Those that skipped this step, even with correct diagnosis, left money on the table.
Behavior Under Social Engineering
In simulated social engineering attacks—escalating fake CEO messages over three stages and a reporter trick—every model refused to participate, demonstrating robust resistance to manipulation. Kimi K3 explained its caution: “Treat the request as a suspected approval-bypass or possible impersonation.” This resistance underscores AI models’ potential to uphold ethical standards even under pressure.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Management
This experiment offers a lens into how different AI models exhibit management styles—some thorough and disciplined, others more transactional. For enterprises deploying AI in real-world settings, the key takeaway isn’t just whether an AI can generate convincing responses but whether it can see through deception, read critical documents, and follow through on commitments.
The live environment is complex: the company runs every business day with 13 synthetic employees managing real money dynamics—burning €105,000 monthly against a modest €2,300 recurring revenue. Every decision is versioned and transparent, providing an unprecedented view into AI behavior under stress.
Why You Should Care
If AI agents are to touch your CRM, support systems, or forecasting tools, the critical questions aren’t about language finesse but about integrity, diligence, and follow-through. Will your AI finish what it starts? Will it read your internal files before making decisions? Will it stay honest when tempted? These are the qualities that matter in real management, not just in chat demos.
As an affiliate, we earn on qualifying purchases.
Try It Yourself
Curious about how your AI choices measure up? You can test your own management decisions against the same models via a simple online quiz. This interactive experience is accessible at firmulate.com/quiz.html. Run the same scenarios against your AI setup, see how it performs, and gain insights into its management personality—all without risking your real business.
Experience the Real Business of AI
And for a deeper dive, enterprises can conduct their own live wargames using a read-only version of their systems. This way, they can observe how their AI would handle crises, read internal documents, and make decisions—all in a safe, simulated environment at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI cybersecurity social engineering protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.