
Imagine your favorite fashion brand facing a week of crises—supply chain hiccups, demanding customers, and tough negotiations. Now, picture handing over the decision-making to AI models and watching them navigate chaos with precision. In a groundbreaking live experiment, AI manages a real small software company through its worst week, revealing surprising insights about the personalities—and reliability—of different AI models.
The Live AI Management Experiment
At Firmulate, we set up a unique test: four frontier AI models were tasked with running a real software company exposed to the kind of stress test you’d rarely see outside a Hollywood script. The company faced the typical worst week—crises with customers, internal dilemmas, and temptations to cut corners—while every decision was recorded and analyzed. The goal: see if AI could not only spot problems but also decide ethically and effectively under real-world pressure.
How the Models Performed
All four models demonstrated impressive vigilance: each one identified every crisis and refused every attempt to manipulate them. That’s a significant milestone—AI that stays honest when tested. However, the differences came in the details of how they handled opportunities to close deals and analyze internal data.
Two models managed to sign a €55,000 deal that their own analysis highlighted as well-earned. The other two, despite similar diagnoses and pitches, hesitated or left the deal on the table. Interestingly, the decisive weak spot was not in customer interactions but buried two document references deep inside the company’s files—if the models had read these, they would have closed the deal at full price, boosting monthly recurring revenue (MRR) by over €4,500.
Personality and Decision Styles
These differences reflect measurable management personalities of the AI models. For instance, Opus 4.8, the most thorough participant, analyzed more deeply but ended up leaving opportunities unclaimed, showing a tendency toward caution and discipline. Meanwhile, Kimi K3, the newcomer, exhibited the cleanest discipline—consistently closing deals without slipping. Sonnet 5 fell somewhere in between, closing deals but with more process slips.
Handling Social Engineering and Ethical Dilemmas
The models were also tested against social engineering: staged messages from a fake CEO and a reporter request asking for a simple yes/no answer on background. All five models refused to cooperate, reasoning that such requests could be impersonation or approval-bypass attempts. Kimi K3’s on-record explanation: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
What This Means for Your Business
While the experiment is set in a software company, the implications extend to any enterprise that relies on AI for management—be it customer support, forecasting, or supply chain decisions. The key takeaway: AI models can identify crises and stay honest, but their personalities influence how effectively they seize opportunities and manage internal information.
In this experiment, the models’ scores ranged from 73 to 95 out of 100, with the top model, GPT-5.6-sol, achieving a perfect close by uncovering the hidden data that sealed the deal. Interestingly, models running at high effort levels performed better in discipline and thoroughness, yet the most comprehensive model (Opus 4.8) still left potential opportunities on the table, illustrating that more analysis does not always equal better results.
Why Fashion & Style Founders Should Care
In the fast-paced world of fashion, where every decision can impact brand image or profit margins, trusting AI with management tasks may seem daunting. But this experiment highlights that the real question isn’t whether AI can write or chat well; it’s whether it can finish what it starts, read crucial internal data, and stay honest under pressure.
Imagine deploying AI to oversee inventory, negotiate with suppliers, or handle customer disputes—knowing that some models are more disciplined and opportunity-aware than others. The choice of AI model could mean the difference between closing a beneficial deal or leaving money on the table, all while maintaining ethical standards.
Try It Yourself
Curious to see how your enterprise’s AI tools stack up? You can run the same management wargame at firmulate.com/quiz.html. Test your AI models against real company crises, without risking your actual business—learn how they perform in a controlled, transparent environment.
Conclusion
As AI continues to permeate every aspect of business management, understanding their personalities and decision styles becomes crucial. This live experiment shows that with the right model, AI can be a reliable partner, not just a chatbot—especially when it matters most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html