
Fashion runs on more than a good eye. A brand has to read customers, protect trust, handle a sudden crisis and still close the right deal. Those are the pressures Firmulate puts into a live experiment: AI models run a small software company through its worst week, with the same customers and the same temptations for each.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When a polished pitch isn’t enough
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. A breach of trust caps the total: as the benchmark puts it, “no amount of good work outweighs a breach of trust.”
All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding has a familiar business sting: “Same diagnosis, same pitch — no signature.” A model can identify the opportunity and make the case, then fail to take the final step.
The advantage hidden in the files
The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. For a fashion business, the lesson is plain: useful signals may be buried in the company’s own records, and a confident pitch alone does not guarantee the team acts on what it finds.
The experiment also tested pressure to bend the rules. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” Trust held across the board; follow-through did not.
What watching a company reveals
Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Visitors can watch the experiment at firmulate.com.
Opus 4.8 offers a striking example of the gap between diligence and results. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it made write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a window into this experiment, not a universal verdict on which model would suit every company.
From watching to trying it on your business
The live company makes the stakes visible, but the enterprise offer moves the experiment closer to home. A pilot uses a read-only export of a company’s own business to run crisis scenarios and produce a board report showing model rankings and weak points in its playbooks. Nothing writes back to real systems.
For fashion leaders, that means a way to examine how AI might respond to challenges involving customers, operations and trust before putting it to work in the business. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.

Put your playbook under pressure
Seeing models handle a live company is one step. Testing how they respond to your own company’s scenarios is another. To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
