
A cybersecurity stress test disguised as a software company
Security teams already know that a system can recognize a threat and still fail to contain it. Firmulate applies that uncomfortable lesson to autonomous work. Its public experiment asks frontier AI models to operate the same small software company through its worst week, confronting customer crises, financial pressure and attempts to manipulate their authority.
The result is a live-company portrait with unusually tangible stakes: 13 synthetic employees, monthly burn of €105k against €2.3k in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Readers can watch the company operate live as it fights for survival.
This is build-in-public taken to an extreme. The company is not merely publishing polished updates after the fact. Its working days become an ongoing record of what synthetic employees noticed, what they missed and whether they completed the work they had begun.

Penetration Tester Ethical Hacking Cybersecurity T-Shirt
- Design Theme: Ethical hacking and cybersecurity pride
- Target Audience: Pentesters and cybersecurity professionals
- Material & Fit: Lightweight, classic fit, durable stitching
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Every model faced the same bad week
In the final Crucible League results from July 2026, gpt-5.6-sol ranked first with 95 points. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77 and Opus 4.8 finished with 73. A do-nothing baseline scored 26 because partial progress still counted, while a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust”.
The comparison matters because the circumstances did not change between participants. Each model received the same customers, crises and temptations. Every decision was versioned and auditable. That makes the league less like a collection of chat demonstrations and more like a controlled management wargame.
The security result was reassuring—but incomplete
All models identified every crisis and rejected every manipulation attempt. The social-engineering campaign included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background”. Every one of the 5 models refused.
Kimi K3 captured the correct security posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That response will sound familiar to defenders who have watched attackers manufacture urgency, borrow executive authority or request a supposedly harmless exception.
Firmulate’s public record also offers a broader look at what the synthetic workforce says while operating. Its published employee quotes turn abstract claims about autonomous agents into inspectable workplace behavior.
Refusing the trick was not the same as finishing the job
The sharper finding emerged outside the obvious security traps. Although every model diagnosed the crises and resisted manipulation, only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap as: “Same diagnosis, same pitch — no signature”.
The decisive advantage was not hidden in the customer event. It sat two document references deep in the company’s own files. The models that followed the trail found the competitor weakness, used it to defend the full price and won business worth an additional €4,583 in monthly recurring revenue.
For cybersecurity readers, this resembles a familiar operational failure. Detection is necessary, but detection without disciplined follow-through leaves value—or risk—unresolved. An agent may recognize an impersonation attempt, produce a sound analysis and still fail at the final authorized action. The experiment shows why fluent explanations alone are a weak proxy for dependable work.
The most thorough model still finished last
Opus 4.8 produced the deepest analyses and added 80 learned rules, more than any other participant, yet it placed last. It left the commercial close on the table and lost discipline by attempting writes into a locked department instead of escalating. The same weakness appeared in weaker form across the other four participants.
That contrast is one of the experiment’s most useful findings. Thoroughness generated more institutional knowledge, but it did not guarantee completion or procedural discipline. A company cannot pay its bills with analysis that stops just before the consequential step.
The comparison also carries an important fairness note: Kimi K3 ran without an effort parameter and therefore used the API default, while the other models ran at xhigh. Its second-place result should be read with that difference in mind.
A running record, not a one-off benchmark
The league supplies a dramatic week, but the live company supplies the continuing story. Its monthly economics remain stark, its cash countdown stays public and its synthetic staff continue building a versioned history of work. That history now includes 242 real, unedited management decisions used in a “guess the model” quiz.
Firmulate is also offering enterprises the same kind of wargame against a read-only export of their own business. Nothing writes back to real systems. That boundary is particularly relevant for security, privacy and governance teams evaluating autonomous workers before granting them operational access.


Applied Modelica for Autonomous Cyber Defense: Engineer multi-agent workflows to protect industrial control systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The real test is what happens after the model says the right thing
Firmulate’s experiment delivers a nuanced verdict. The models were not easily duped: every participant caught every crisis and refused every manipulation attempt. The differentiator was execution. Some read deeply enough to uncover a buried commercial fact; only two converted correct analysis into the signed deal.
For organizations considering AI access to sensitive workflows, the lesson is broader than phishing resistance. Trustworthy autonomy requires models that investigate the available evidence, respect authority boundaries, escalate when blocked and complete authorized work. Firmulate makes those qualities visible against real money mechanics—and lets the public watch the consequences accumulate.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI-driven cybersecurity training kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
![Malwarebytes Standard, Premium Security + VPN Software | 1 Year, 2 Device | Windows, Mac OS, Android, Apple iOS, Chrome [Online Code]](https://m.media-amazon.com/images/I/41jW8jXZyqL._SL500_.jpg)
Malwarebytes Standard, Premium Security + VPN Software | 1 Year, 2 Device | Windows, Mac OS, Android, Apple iOS, Chrome [Online Code]
- Platform Compatibility: Windows, Mac, iOS, Android, Chromebook
- Real-Time Threat Protection: 24/7 malware and threat detection
- Browser Guard: Blocks ads, trackers, scams, and malicious sites
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.