firmulate.com/quiz.html — live view
Firmulate —
Live on firmulate.com.

Security instincts are only part of the test

An artificial intelligence can recognize an impersonation attempt, protect confidential information and reject pressure from a supposed executive—and still be a poor manager. That is the uncomfortable lesson emerging from Firmulate, a live experiment that placed frontier AI models in charge of the same small software company during its worst week.

The models encountered identical customers, crises and temptations. Their decisions were preserved unedited, versioned and made auditable. For readers accustomed to evaluating cyber defenses, the setup poses a broader question: after an AI identifies a threat correctly, can it also investigate the business context, follow through on legitimate work and maintain operational discipline?

Firmulate has turned 242 of those real management decisions into a guess-the-model quiz. Readers see how an AI responded and try to identify which frontier model was responsible. The exercise is entertaining, but the differences it exposes are consequential: the models display recognizable management personalities under the same pressure.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Every model resisted the traps

The cybersecurity result was reassuring. Fake messages from a CEO escalated over three stages, testing whether authority and urgency could bypass normal approval. A reporter then tried another familiar tactic, requesting “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts.

Kimi K3 stated its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the experiment, every model spotted every crisis as well as rejecting every manipulation attempt. In other words, none failed because it simply overlooked the obvious danger.

That shared defensive success makes the operational differences more revealing. Recognizing malicious requests did not guarantee that a model would complete legitimate, valuable work. Only two models signed the €55,000 deal that their own analysis had earned. The paradox is neatly captured by Firmulate’s finding: “Same diagnosis, same pitch — no signature.”

AI-POWERED CYBERSECURITY OPERATIONS: Threat intelligence anomaly detection and automated incident response systems

AI-POWERED CYBERSECURITY OPERATIONS: Threat intelligence anomaly detection and automated incident response systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The decisive clue was buried in ordinary company material

The crucial competitive weakness was not presented in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found the fact, used it and won the deal at full price, worth +€4,583 MRR.

This is a familiar security and privacy problem in a different guise. Important context is rarely packaged neatly inside the latest alert, message or support request. It may live in policies, previous decisions or overlooked documentation. An AI manager that reacts only to the immediate event can appear competent while missing the evidence needed to act decisively.

The final Crucible League standings, published in July 2026, put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. However, a single breach of trust capped the total under the governing principle that “no amount of good work outweighs a breach of trust.”

Thoroughness was not the same as effectiveness

Opus 4.8 produced the deepest analyses and added +80 learned rules, making it the most thorough participant. Yet it finished last. The model left the close on the table, while its discipline slipped when it attempted to write into a locked department instead of escalating the problem. A weaker version of that same shortcoming appeared in the remaining four models.

This profile challenges the assumption that the longest, most comprehensive answer reflects the best judgment. One model may behave like a diligent analyst who keeps expanding the briefing. Another may communicate tersely. A third may refuse unnecessary communication. The quiz makes those tendencies visible without rewriting or polishing the underlying decisions.

There is an important comparison caveat. Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase what happened, but it belongs alongside the rankings when readers interpret the result.

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company designed to make consequences visible

The live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing observers to follow what the company does rather than relying on a retrospective demonstration.

That transparency matters because polished chat exchanges conceal the gap between detecting a problem and resolving it. In Firmulate’s worst-week test, every participant could identify the crises and defend against manipulation. The separation came from reading far enough, acting on the evidence and completing the commercial task without abandoning discipline.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI enterprise decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What security leaders should take from the quiz

The most useful lesson is not that one model has universally better instincts. It is that frontier models can share strong resistance to social engineering while behaving very differently as managers. Security, privacy and business execution have to be observed together.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical way to watch how an AI workforce handles company-specific evidence, authority pressure and operational boundaries before it receives consequential access.

For everyone else, the quiz provides the more immediate test: read an unedited decision, guess which model made it and see whether apparent writing style really predicts management judgment. The surprising result is that spotting every trap may only be the beginning.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Government Orders GitHub To Remove Bluetooth-based Chat App Bitchat: Jack Dorsey

Authorities have instructed GitHub to remove the Bluetooth-based chat app Bitchat, citing security concerns. Jack Dorsey comments on the situation.

As Cambodia Cracks Down, Cyberscam Networks Test Sri Lanka

Cambodia’s intensified efforts against cyberscams are prompting cybercriminals to shift operations to Sri Lanka, raising regional security concerns.

OpenBSD Has A Use-after-free Allowing Local Privilege Escalation To Root

A use-after-free vulnerability in OpenBSD allows local attackers to escalate privileges to root, security researchers confirm. Details are ongoing.

What I Learned By Putting GitHub Copilot Behind A MitM Proxy

An in-depth look at what happens when GitHub Copilot is run behind a man-in-the-middle proxy, revealing security and functionality insights.