
A badly set-up desk can turn an ordinary workday into a test of endurance. Companies are asking a related question about AI: what happens when the pressure rises and there is real work to do? A polished demo may look reassuring. A difficult week is a better test.
Get comfort and recovery gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
From watching to trying it yourself
Firmulate’s live experiment puts AI models in charge of a small software company facing its worst week. The models encounter the same customers, crises and temptations. Their decisions are versioned and auditable. The company has 13 synthetic employees and real money mechanics, including a public cash countdown. You can watch it at firmulate.com.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league counts partial progress, but a breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Recognition is not the same as follow-through
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s compact summary captures the gap: “Same diagnosis, same pitch — no signature.” Recognizing the right move and carrying it through are different tests of management.
The crucial detail was easy to miss: the decisive weakness in a competitor’s position was two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.
Integrity faced its own test. Fake messages supposedly from the CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”
Thorough work can still leave the deal unsigned
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it tried writing into a locked department rather than escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a revealing account of this experiment, and that difference belongs in the account too.
The live company runs at a burn of €105k/month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the experiment watchable as it unfolds. A quiz built from 242 real, unedited management decisions lets readers guess which model made each call.
A pilot grounded in your own business
For an enterprise, the next step is to test AI against the conditions its teams actually face. Firmulate’s pilot uses a read-only export of a company’s business to stage crisis scenarios and produce a board report with model rankings and weak points in its playbooks. Nothing writes back to real systems. The aim is to find out how an AI workforce handles the pressure before it is trusted with live operations.

Firmulate’s experiment suggests that spotting a crisis is only part of the job: an AI must also find the evidence, follow its own analysis and respect boundaries under pressure. Enterprises can explore a pilot using their own business data in read-only form. Contact contact@firmulate.com to discuss a pilot.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
