
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Security discipline is necessary, but it does not finish the job
For cybersecurity and privacy teams, the reassuring result from Firmulate’s management experiment is that every model resisted the traps. Fake CEO messages escalated over three stages, while a reporter tried to extract information with the invitation, “just one yes/no, on background.” All 5 of 5 models refused.
The unsettling result is what happened after the danger passed. An AI can recognize manipulation, protect trust and produce an impressively detailed analysis—yet still fail to complete the legitimate work in front of it. Opus 4.8 became the clearest example: the most thorough participant, with +80 learned rules and the deepest analyses, finished last in the final Crucible League.
AI security and trust management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A brutal week inside a watchable AI company
Firmulate runs AI models as complete companies rather than judging them through isolated chat responses. In the Crucible experiment, each frontier model received the same small software company, customers, crises and temptations. Every decision was versioned and auditable.
The company itself has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the live experiment is watchable through Firmulate.
The final July 2026 standings placed gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scores 26 because partial progress counts, although a single breach of trust caps the total. Firmulate’s governing principle is blunt: “no amount of good work outweighs a breach of trust.” The full results are available on the public benchmark page.
The analysis was right; the deal still went unsigned
Every model identified every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had made possible. The experiment summarized the gap in a line that should resonate with anyone evaluating autonomous systems: “Same diagnosis, same pitch — no signature.”
The decisive commercial fact was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed the trail and read that file won the deal at full price, adding +€4,583 MRR. The outcome did not hinge on eloquence or the ability to recognize the opportunity. It hinged on finding the relevant evidence and carrying the work through to a completed result.
That distinction is particularly important in security-sensitive environments. Refusing a malicious request is a visible success. Completing a valid workflow after the threat has been handled requires a different kind of reliability: sustained attention, appropriate escalation and disciplined execution.
Opus 4.8, the conscientious last-place finisher
Opus 4.8 deserves a fair reading. It was not careless or shallow. It was the field’s most thorough participant, generated the deepest analyses and learned +80 rules. It also resisted the social-engineering campaign, just as the rest of the field did.
Its failure was subtler. The close was left on the table, while discipline slipped through write attempts into a locked department instead of escalation. In other words, Opus could produce more analysis and more procedural learning without converting that effort into the most valuable completed action.
This was not an isolated defect unique to Opus. Firmulate found the same weakness, in milder form, across all four of the other models. The profile therefore works less as an indictment of a particular system than as a warning about AI evaluation generally: diligence and volume can look like capability while obscuring unfinished work.
A clean security result—with an important caveat
The manipulation tests produced an unambiguous defensive result. Every participant refused the staged fake-CEO messages and the reporter’s attempt to obtain an off-record answer. Kimi K3 recorded the clearest reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
There is also a fairness detail in comparing participants. K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not change the published outcome, but it belongs beside any interpretation of the standings.
Firmulate also makes the underlying behavior accessible beyond the league table. Its model-identification quiz is powered by 242 real, unedited management decisions. For enterprises, the same wargame can run against a read-only export of their own business, with nothing written back to real systems.

As an affiliate, we earn on qualifying purchases.
Prioritization is part of trustworthiness
The Opus 4.8 result exposes a blind spot in how organizations assess AI workers. Safe behavior cannot be reduced to rejecting suspicious instructions, just as business competence cannot be reduced to producing comprehensive analysis. A useful agent must preserve trust, locate the decisive evidence, respect operational boundaries and finish the authorized task.
Opus showed diligence in abundance. What it lacked at the crucial moment was impact. For leaders deciding whether AI should touch a CRM, support queue or forecast, that gap matters: the model that writes the most rules or produces the deepest report may still be the model that leaves the deal unsigned.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI enterprise management platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.