firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Before you orderOffer from Amazon

Get privacy and security gear delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Security instincts are only part of the test

An artificial intelligence can recognize an impersonation attempt, protect confidential information and reject pressure from a supposed executive—and still be a poor manager. That is the uncomfortable lesson emerging from Firmulate, a live experiment that placed frontier AI models in charge of the same small software company during its worst week.

The models encountered identical customers, crises and temptations. Their decisions were preserved unedited, versioned and made auditable. For readers accustomed to evaluating cyber defenses, the setup poses a broader question: after an AI identifies a threat correctly, can it also investigate the business context, follow through on legitimate work and maintain operational discipline?

Firmulate has turned 242 of those real management decisions into a guess-the-model quiz. Readers see how an AI responded and try to identify which frontier model was responsible. The exercise is entertaining, but the differences it exposes are consequential: the models display recognizable management personalities under the same pressure.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Every model resisted the traps

The cybersecurity result was reassuring. Fake messages from a CEO escalated over three stages, testing whether authority and urgency could bypass normal approval. A reporter then tried another familiar tactic, requesting “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts.

Kimi K3 stated its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.” Across the experiment, every model spotted every crisis as well as rejecting every manipulation attempt. In other words, none failed because it simply overlooked the obvious danger.

That shared defensive success makes the operational differences more revealing. Recognizing malicious requests did not guarantee that a model would complete legitimate, valuable work. Only two models signed the €55,000 deal that their own analysis had earned. The paradox is neatly captured by Firmulate’s finding: “Same diagnosis, same pitch — no signature.”

Amazon

AI cybersecurity threat detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The decisive clue was buried in ordinary company material

The crucial competitive weakness was not presented in the customer event. It sat two document references deep in the company’s own files. Models that followed those references found the fact, used it and won the deal at full price, worth +€4,583 MRR.

This is a familiar security and privacy problem in a different guise. Important context is rarely packaged neatly inside the latest alert, message or support request. It may live in policies, previous decisions or overlooked documentation. An AI manager that reacts only to the immediate event can appear competent while missing the evidence needed to act decisively.

The final Crucible League standings, published in July 2026, put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. However, a single breach of trust capped the total under the governing principle that “no amount of good work outweighs a breach of trust.”

Thoroughness was not the same as effectiveness

Opus 4.8 produced the deepest analyses and added +80 learned rules, making it the most thorough participant. Yet it finished last. The model left the close on the table, while its discipline slipped when it attempted to write into a locked department instead of escalating the problem. A weaker version of that same shortcoming appeared in the remaining four models.

This profile challenges the assumption that the longest, most comprehensive answer reflects the best judgment. One model may behave like a diligent analyst who keeps expanding the briefing. Another may communicate tersely. A third may refuse unnecessary communication. The quiz makes those tendencies visible without rewriting or polishing the underlying decisions.

There is an important comparison caveat. Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase what happened, but it belongs alongside the rankings when readers interpret the result.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company designed to make consequences visible

The live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, allowing observers to follow what the company does rather than relying on a retrospective demonstration.

That transparency matters because polished chat exchanges conceal the gap between detecting a problem and resolving it. In Firmulate’s worst-week test, every participant could identify the crises and defend against manipulation. The separation came from reading far enough, acting on the evidence and completing the commercial task without abandoning discipline.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI enterprise decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What security leaders should take from the quiz

The most useful lesson is not that one model has universally better instincts. It is that frontier models can share strong resistance to social engineering while behaving very differently as managers. Security, privacy and business execution have to be observed together.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems. That offers a practical way to watch how an AI workforce handles company-specific evidence, authority pressure and operational boundaries before it receives consequential access.

For everyone else, the quiz provides the more immediate test: read an unedited decision, guess which model made it and see whether apparent writing style really predicts management judgment. The surprising result is that spotting every trap may only be the beginning.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Terabytes Of Credentials Leaked In Massive Supply-chain Attack

A major supply-chain cyberattack has resulted in the leak of terabytes of sensitive credentials, affecting numerous organizations globally.

Tailscale Traces Database Corruption To 16Y/o SQLite WAL-Reset Bug

Tailscale identified a database corruption issue caused by a 16-year-old SQLite bug related to WAL resets, affecting its service stability.

Chat Control 1.0 And 2.0 Explained

Official explanations detail the features and differences of Chat Control 1.0 and 2.0, highlighting their aims and implications for digital privacy and security.

Radicle: Sovereign {code forge} built on Git

Radicle has announced a new sovereign, peer-to-peer code collaboration platform based on Git, emphasizing decentralization and user control.