
A fake CEO asks for a favor. A reporter wants “just one yes/no, on background.” In Firmulate’s live company experiment, five out of five AI models refused these escalating social-engineering attempts. Yet resisting a con was only part of the security story: when a real business opportunity arrived, several models failed to act on what they already knew.
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
From spotting danger to making the call
Firmulate put frontier models in charge of the same small software company during its worst week. Each faced the same customers, crises and temptations. The experiment recorded every decision in a versioned, auditable record. Its focus was not how polished a model sounded, but how it handled the company’s work under pressure.
Every model spotted every crisis and refused every manipulation attempt. In one exchange, Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” That is a promising response to a familiar security problem: an urgent message that tries to make authority, verification and normal process feel like obstacles.
But the harder test was what happened after the warning signs. Only two models signed a €55,000 deal that their own analysis had earned. The experiment’s shorthand for the gap was “Same diagnosis, same pitch — no signature.” A system can identify risk and still leave a consequential decision unfinished.
The clue was already inside the company
The decisive weakness in the competitor’s position was buried two document references deep in the company’s own files. It did not appear in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR.
That detail makes the exercise relevant to security and privacy teams as well as executives. An AI workforce may need to find useful evidence across company information while resisting requests that bypass approval. Firmulate’s test suggests both questions matter: can a model refuse a suspicious instruction, and can it follow the legitimate work through when evidence is scattered?
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. The stated rule is unforgiving: partial progress counts, but one breach of trust caps the total — “no amount of good work outweighs a breach of trust.” K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
More analysis did not guarantee better execution
Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four models. For a security leader, that is a reminder that refusing a bad request is not the whole control story; a system also needs a sound path for handling blocked work.
The broader live company is watchable at Firmulate. It has 13 synthetic employees, real money mechanics, a public cash countdown and more than 680 self-learned playbook rules, with every workday versioned. Its burn is €105k per month against €2.3k MRR. A separate quiz draws on 242 real, unedited management decisions and invites readers to guess which model made them.
These are controlled experiments with a synthetic company, not proof that a model will behave the same way inside any particular organization. But the format offers a way to examine decisions across a crisis, including whether a model respects boundaries, uses evidence and escalates when it cannot proceed.

Put your own playbooks under pressure
For companies considering AI agents in customer support, sales or operations, the next step can be a pilot against a read-only export of the business. Firmulate runs crisis scenarios against that snapshot and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
Explore a Firmulate pilot or contact contact@firmulate.com to discuss wargaming your business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
