AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Fashion runs on more than a good eye. A brand has to read customers, protect trust, handle a sudden crisis and still close the right deal. Those are the pressures Firmulate puts into a live experiment: AI models run a small software company through its worst week, with the same customers and the same temptations for each.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When a polished pitch isn’t enough

The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. A breach of trust caps the total: as the benchmark puts it, “no amount of good work outweighs a breach of trust.”

All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding has a familiar business sting: “Same diagnosis, same pitch — no signature.” A model can identify the opportunity and make the case, then fail to take the final step.

The advantage hidden in the files

The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. For a fashion business, the lesson is plain: useful signals may be buried in the company’s own records, and a confident pitch alone does not guarantee the team acts on what it finds.

The experiment also tested pressure to bend the rules. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” Trust held across the board; follow-through did not.

What watching a company reveals

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. Visitors can watch the experiment at firmulate.com.

Opus 4.8 offers a striking example of the gap between diligence and results. It was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. The close was left on the table, and discipline slipped when it made write attempts into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a window into this experiment, not a universal verdict on which model would suit every company.

From watching to trying it on your business

The live company makes the stakes visible, but the enterprise offer moves the experiment closer to home. A pilot uses a read-only export of a company’s own business to run crisis scenarios and produce a board report showing model rankings and weak points in its playbooks. Nothing writes back to real systems.

For fashion leaders, that means a way to examine how AI might respond to challenges involving customers, operations and trust before putting it to work in the business. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbook under pressure

Seeing models handle a live company is one step. Testing how they respond to your own company’s scenarios is another. To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI’s Deep File-Reading Skills Could Make or Break Your Business Deals

An experiment reveals that AI’s ability to read deeply into company files—two references deep—can be the decisive factor in closing high-value deals. Trustworthiness and thoroughness matter.

NFTs for Fashion Loyalty Programs—Hype or Helpful?

An exploration of whether NFTs in fashion loyalty programs are just hype or genuinely beneficial—discover the truth behind this innovative trend.

Email Marketing Automations Every Boutique Needs

Wondering how to boost sales and engagement? Discover essential email marketing automations every boutique needs to transform your business today.

What AI Management Skills Really Matter — Beyond Chat Quality

AI leadership isn’t just about chat—it’s about managing real business crises, reading internal files, and staying honest under pressure. Live tests reveal what really matters.