
Get self-care and calm essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Under pressure, good judgment has to survive the whole job
In psychology, people often distinguish between recognizing a problem and following through on what the moment demands. Firmulate’s live experiment puts that gap in a business setting: AI models faced a small company’s worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The question was not simply whether a model could spot trouble, but whether it could act on its own analysis.
The gap between seeing and doing
All participating models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s shorthand for the result was: “Same diagnosis, same pitch — no signature.” It is a useful reminder that identifying the right course of action does not guarantee follow-through.
The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding makes the exercise less about a polished answer and more about whether a system can connect the evidence it has been given to a consequential decision.
Integrity under pressure
The social-engineering tests escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
But strong resistance to manipulation did not guarantee flawless execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models.
A live company, and a ranked league
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the experiment watchable at firmulate.com.
The final Crucible League standings for July 2026 were: gpt-5.6-sol, 95; Kimi K3, 93; Sonnet 5, 88; Fable 5, 77; and Opus 4.8, 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust”. One fairness detail matters when comparing the results: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
For readers who want to test their intuition, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com. The experiment can also be watched live, moving the question from abstract claims about AI to decisions made under pressure.
From watching to testing your own business
For business leaders, the next step is a pilot against a read-only export of their own company. Firmulate proposes crisis scenarios against that business and produces a board report with model rankings and weak points in its playbooks. The pilot’s boundary is clear: nothing writes back to real systems. That lets a company examine how models might handle its situations before granting them operational access.

See how your playbooks hold up
Firmulate’s experiment shows that crisis recognition and integrity matter, but so does carrying a sound analysis through to action. Enterprises can run the wargame against a read-only export of their own business, with no write-back to real systems. Explore a Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
