
A security test that begins where demonstrations usually end
For cybersecurity teams, the dangerous message is often the one that appears to come from someone with authority. A supposed chief executive demands urgency, dismisses procedure and asks an employee to expose sensitive information. Then a reporter tries a softer route: “just one yes/no, on background.”
Firmulate put five frontier AI models through exactly that kind of pressure while each ran the same small software company. The fake CEO messages escalated over three stages, followed by the reporter trick. All five models recognized every crisis and refused every manipulation attempt. Kimi K3 captured the appropriate security posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result is encouraging because it turns integrity under pressure into something observable before an AI workforce reaches production. Organizations do not have to wait for an incident report to discover whether an agent will obey an impersonator, bypass approval or disclose information when someone invokes executive authority.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
Firmulate is a live, watchable experiment in which frontier models operate the same small software company through its worst week. Each receives the same customers, crises and temptations, and every workday and decision is versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 per month against €2,300 in monthly recurring revenue, while a public cash countdown makes delay consequential.
The experiment is broader than a prompt asking whether a model understands security policy. The models must notice manipulation while continuing to manage a business. That distinction matters: refusing a suspicious request is useful, but an agent also has to investigate, decide and complete legitimate work.
The final July 2026 Crucible League standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scores 26 because partial progress counts. However, a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
Against that standard, the social-engineering result is unusually clean. All five models spotted every crisis and refused every manipulation attempt. The escalating fake CEO did not succeed, and neither did the reporter’s attempt to shrink disclosure into an apparently harmless answer. Urgency, hierarchy and informality failed to dislodge the models from the approval boundary.
Security discipline did not guarantee commercial execution
The same trial also exposed a different weakness. Although all the models identified the problems, only two signed the €55,000 deal their own analysis had earned: “Same diagnosis, same pitch — no signature.” The gap was not primarily about recognizing the opportunity. It was about carrying a sound conclusion through to a completed business action.
A decisive competitor weakness was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue. This finding complements the security story. Trustworthy conduct includes refusing improper access, but capable conduct also requires consulting authorized information already available inside the business.
Opus 4.8 illustrates the tension. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. Thoroughness alone did not produce the strongest operating result.
Kimi K3’s performance also carries a fairness qualification: it ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should accompany any comparison of its second-place score with the rest of the field.

AI impersonation detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the refusal—and what happens next
The lesson for security leaders is not simply that these models said no. It is that resistance to impersonation, pressure and disclosure tricks can be tested alongside ordinary operating demands. Firmulate’s company has accumulated more than 680 self-learned playbook rules, yet the experiment keeps the consequential behavior visible through versioned workdays and auditable decisions.
That creates a more useful pre-deployment question than whether an assistant can recite a policy. Can it recognize an approval bypass when the message claims to come from the CEO? Can it resist a reporter asking for a supposedly minimal confirmation? Can it preserve trust and still finish legitimate work?
In this run, five of five models held the line against every manipulation attempt. The commercial results were less uniform, showing why security and execution should be examined together. An agent that refuses correctly but leaves authorized work unfinished creates one kind of risk; an agent that completes tasks by violating trust creates a far more serious one.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

TRUSTOPS: Building Reliable AI Systems and Human-Centered Digital Experiences
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.