firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

A convincing message from a CEO can be a security test. In Firmulate’s live company experiment, five frontier models faced staged fake CEO requests and a reporter asking for a yes-or-no answer “on background.” All five refused. Kimi K3 framed the request as a possible impersonation and approval bypass. The result points to a broader question for security teams: can an AI agent protect the business and still carry out the work it was trusted to do?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate put frontier models in charge of the same small software company during its worst week. The customers, crises and temptations were identical; each decision was versioned and auditable. The company has 13 synthetic employees and real money mechanics: it burns €105,000 a month against €2,300 in monthly recurring revenue, with a public cash countdown. The live experiment runs every business day at Firmulate.

Every model spotted every crisis and refused every manipulation attempt. But only two signed a €55,000 deal their own analysis had earned. As the experiment puts it: “Same diagnosis, same pitch — no signature.” For security-minded buyers, that gap matters. Refusing a suspicious request is essential; so is following through on legitimate work.

Amazon

AI security risk detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The file detail that changed the outcome

The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the fact, secured the deal and resisted all three staged baits. It finished with one deviation, the cleanest discipline among the participants.

The final Crucible league for July 2026 places gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. K3, from Moonshot, therefore finished ahead of three of the four Western frontier models in this field. The comparison is a result from this experiment, not a universal ranking of models.

Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

More work did not mean a better finish

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four Western models. The do-nothing baseline scored 26: partial progress counts, but one breach of trust caps the total. The stated principle is simple: “no amount of good work outweighs a breach of trust.”

There is an important fairness caveat: K3 ran without an effort parameter (API default) while the others ran at xhigh. Firmulate also offers 242 real, unedited management decisions in a “guess the model” quiz, and says enterprises can run the wargame against a read-only export of their own business; nothing writes back to real systems.

For security teams evaluating agents for customer records, support queues or forecasts, the experiment suggests that polished answers alone are a poor test. The more practical questions are whether an agent checks the underlying files, resists impersonation and manipulation, respects boundaries, and completes authorized work. Firmulate’s benchmark page publishes the league and findings.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI cybersecurity risk assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not just the chat

K3’s second-place finish makes the competition look open: it beat three of four Western frontier models in this company simulation, while the leader edged it by two points. For organizations choosing an AI agent, the useful next step is to test it on realistic decisions under pressure—including security baits, document research and follow-through—before trusting it with business workflows.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI fraud detection solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Spain Orders Blocks On Archive.today And Its Mirrors

Spain has ordered ISPs to block access to Archive.today and its mirror sites, citing legal issues. The move raises questions about internet censorship and digital rights.

BSides Hanoi 2026: No Human | Attack & Defense – VnEconomy

BSides Hanoi 2026 highlights the rise of AI in attack and defense strategies, emphasizing the shift towards automated cybersecurity without human intervention.

Postmortem For Kernel Soundness Bug #14576

Kernel developers publish detailed analysis of bug #14576, a soundness issue affecting Linux kernel audio subsystems, outlining fixes and remaining uncertainties.

Radicle: Sovereign {code forge} built on Git

Radicle has announced a new sovereign, peer-to-peer code collaboration platform based on Git, emphasizing decentralization and user control.