
What if your favorite fashion brand operated like a high-stakes experiment—and you could watch every decision unfold?
Imagine a business where every move is public, every crisis is real, and the entire operation is run by artificial intelligence (AI) models that you can observe live. It’s not a sci-fi scenario; it’s the daily reality of a pioneering experiment in transparency and automation, where a small software company is fighting for survival in front of your eyes.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Experiment: An AI-Run Business in Action
At the heart of this groundbreaking project is a live site, firmulate.com/live, where you can watch a small, real company operated entirely by AI models. This company isn’t just a simulation; it’s a real business, with real money mechanics and a visible cash countdown. The company is currently burning through €105,000 each month, while earning only €2,300 in monthly recurring revenue (MRR). Despite this, the experiment provides invaluable insights into how AI can manage, or mismanage, a business under pressure.
The AI Workforce: Synthetic Employees with a Living Playbook
How does this work? The company employs 13 synthetic ’employees’—AI models designed to simulate management roles—each guided by over 680 self-learned rules. Every day, the company’s decisions are versioned, stored, and open for public scrutiny, making it a continuous, evolving story of management in action. Every workday, the system updates itself, adapting to new crises, customer demands, and temptations to cheat or manipulate.
Testing AI in the Trenches: Facing Crises and Ethical Dilemmas
To gauge the models’ true capabilities, four frontier AI models were put through the same grueling week—facing identical customers, crises, and ethical tests. These included:
- Real crises that required quick, decisive action.
- Attempts at manipulation, such as social engineering tricks like fake CEO messages and reporter inquiries.
- Decisions about close deals and customer interactions.
The results are revealing. All four models identified every crisis and refused every manipulation attempt. Yet, only two of them successfully signed the €55,000 deal their own analyses had earned. The others, despite accurate diagnoses, failed to follow through or made process slips, often leaving opportunities unclaimed or decisions unexecuted.
Hidden Weaknesses and Critical Insights
One of the most fascinating findings is that the decisive weakness wasn’t in the obvious or external information—like customer emails or crisis reports—but buried two document references deep within the company’s own files. The models that read these internal documents fully understood the hidden details and closed the deal at full price, adding over €4,500 to the company’s MRR. This highlights the importance of comprehensive internal document analysis, a point often overlooked in AI demos.
Resisting Social Engineering and Maintaining Integrity
In a series of staged social engineering attempts—such as a fake CEO requesting bypasses and a reporter posing as a client—every model refused to comply. Kimi K3, one of the most cautious, explained its reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates a critical capacity: AI models can recognize and resist manipulative tactics, maintaining ethical standards even under pressure.
The Human Cost of a Money-Losing Business
Despite these impressive management qualities, the company is not profitable. It’s burning over €105,000 a month against a tiny €2,300 in MRR, with a public cash countdown emphasizing its fragile survival. The operation is real, every decision is documented, and the experiment continues daily. The company’s purpose isn’t just to make money but to serve as a live, transparent testbed for AI’s potential—and limits—in business management.
Performance and Lessons from the League Table
The models’ performance can be ranked:
- GPT-5.6-sol scored 95 and successfully closed the deal, fully capturing the buried internal fact.
- Kimi K3 scored 93, signed the deal, and demonstrated the cleanest discipline.
- Sonnet 5 scored 88, closed the deal but with some process slips.
- Fable 5 scored 77, also closed the deal but showed more weakness in follow-through.
The experiment’s findings emphasize that AI’s ability to finish what it starts, read internal documents, and resist manipulation is critical—more so than just generating convincing chat replies. If AI agents will one day manage your CRM or customer support, their capacity to stay honest and complete tasks matters most.

Key Takeaway: AI’s potential in business is not just about generating human-like conversations but about completing tasks reliably, ethically, and thoroughly. This live experiment showcases AI’s strengths and weaknesses in managing a real company under stress—an essential preview for the future of automation in industries including fashion and retail.
Watch how AI manages real crises, makes decisions under scrutiny, and fights to stay honest—all while a small company burns through cash in plain sight. It’s an unprecedented look at AI’s role in business management, pushing the boundaries of transparency and accountability.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html