firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get self-care and calm essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Under pressure, good judgment has to survive the whole job

In psychology, people often distinguish between recognizing a problem and following through on what the moment demands. Firmulate’s live experiment puts that gap in a business setting: AI models faced a small company’s worst week, with the same customers, crises and temptations. Every decision was versioned and auditable. The question was not simply whether a model could spot trouble, but whether it could act on its own analysis.

The gap between seeing and doing

All participating models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. The experiment’s shorthand for the result was: “Same diagnosis, same pitch — no signature.” It is a useful reminder that identifying the right course of action does not guarantee follow-through.

The deal hinged on a detail buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding makes the exercise less about a polished answer and more about whether a system can connect the evidence it has been given to a consequential decision.

Integrity under pressure

The social-engineering tests escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

But strong resistance to manipulation did not guarantee flawless execution. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four models.

A live company, and a ranked league

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR. Its public cash countdown, 680+ self-learned playbook rules and versioned workdays make the experiment watchable at firmulate.com.

The final Crucible League standings for July 2026 were: gpt-5.6-sol, 95; Kimi K3, 93; Sonnet 5, 88; Fable 5, 77; and Opus 4.8, 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust”. One fairness detail matters when comparing the results: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

For readers who want to test their intuition, 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com. The experiment can also be watched live, moving the question from abstract claims about AI to decisions made under pressure.

From watching to testing your own business

For business leaders, the next step is a pilot against a read-only export of their own company. Firmulate proposes crisis scenarios against that business and produces a board report with model rankings and weak points in its playbooks. The pilot’s boundary is clear: nothing writes back to real systems. That lets a company examine how models might handle its situations before granting them operational access.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how your playbooks hold up

Firmulate’s experiment shows that crisis recognition and integrity matter, but so does carrying a sound analysis through to action. Enterprises can run the wargame against a read-only export of their own business, with no write-back to real systems. Explore a Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

This article is for informational purposes only and is not medical advice. Always consult a qualified healthcare professional about your specific situation.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Procedural SVG Generation: A Look Inside “Alpenpost Studio – Procedural Alpine Travel Posters” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“Alpenpost Studio…

AI Management Tests Reveal Gaps Beyond Chat Quality: The Price of Trust and Discipline in Business Decision-Making

Live AI management tests reveal that performance under pressure, thoroughness, and honesty are vital—bivotal qualities often invisible in chat demos but crucial for real-world business.

Ethical Performance Evaluation: Monitor Your Impact

Learn the key steps for How to Monitor and Evaluate Your Ethical Performance and Impact effectively with our comprehensive guide.