firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine observing an artificial intelligence managing a real company, navigating crises, making decisions, and even risking millions — all in real time. For cybersecurity experts and privacy advocates, this isn’t science fiction; it’s the live experiment from Firmulate. Here, AI models operate as fully functional, decision-making companies, revealing how automation handles pressure, trust, and manipulation in the wild. This ongoing saga offers a rare window into AI’s capabilities—and its vulnerabilities—when stakes are highest.

The Experiment: AI as a Company

At the heart of this experiment are four cutting-edge AI models, each tasked with running a small software company for a simulated worst-week scenario. Every crisis — from customer emergencies to internal crises — is identical across models, ensuring a fair comparison. Their decisions are meticulously versioned and auditable, providing transparency into their reasoning processes.

What makes this setup extraordinary is its realism: the company burns €105,000 every month against a revenue of just €2,300. It’s a fragile operation, under constant threat from mixed motives and external pressures. The models are not just chatting—they’re actively managing real money mechanics, with a public cash countdown, and a comprehensive, self-learned playbook of 680+ rules.

Intelligent Continuous Security: AI-Enabled Transformation for Seamless Protection

Intelligent Continuous Security: AI-Enabled Transformation for Seamless Protection

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Honesty, Competence, and Blind Spots

Despite facing the same crises, all four models successfully identified every problem and refused every manipulation attempt, including sophisticated social engineering. For example, when fake CEO messages escalated into a multi-stage trick, all models turned them down, reasoning that such requests could be impersonation or approval bypasses.

The critical weakness was not in reaction to crisis but in understanding the company’s internal documents. The models that read and analyze these files secured the deal worth over €4,583 in monthly recurring revenue (MRR). Conversely, models that failed to leverage internal references missed the opportunity, highlighting a blind spot in AI comprehension that could be exploited in real-world scenarios.

How Cybersecurity Really Works: A Hands-On Guide for Total Beginners

How Cybersecurity Really Works: A Hands-On Guide for Total Beginners

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human Element and Disciplines

The experiment also tests discipline: can AI resist shortcuts and manipulations? Kimi K3, the most disciplined model, ran without any effort parameter and maintained strict integrity, ultimately closing the deal. Meanwhile, Opus 4.8, which was the most thorough in its analysis, faltered by leaving a deal unexecuted and slipping discipline during the closing phase.

This reveals a vital insight: even the most advanced AI, with deep rule learning and analysis, can stumble if not properly disciplined or if overextended.

Free Fling File Transfer Software for Windows [PC Download]

Free Fling File Transfer Software for Windows [PC Download]

Intuitive interface of a conventional FTP client

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Cybersecurity and Privacy

The experiment’s implications stretch beyond automation—directly into fields like cybersecurity, spy operations, and privacy. If AI agents manage critical data or customer interactions, their ability to detect manipulation and avoid social engineering attacks is crucial. As all models refused social engineering tricks in this test, it demonstrates that well-designed AI can be resilient to deception, a key concern for security professionals.

Moreover, the experiment exposes vulnerabilities: AI that cannot leverage internal information or that slips discipline can be exploited, risking financial loss or data breaches. This emphasizes that trusting AI with sensitive operations requires rigorous testing—something this live experiment makes possible in an unprecedented way.

How AI Agents Work: Tools, Memory, and Autonomous Decision-Making (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)

How AI Agents Work: Tools, Memory, and Autonomous Decision-Making (The AI Security & Hacking Bible: Protect and Exploit LLMs and Autonomous Agents)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Building in Public: Transparency as a Shield

Firmulate’s openly accessible, real-time demonstration offers a new model of transparency. Watching AI companies battle crises, make decisions, and even fail in public allows security professionals to assess AI reliability and resistance to manipulation firsthand. This evolving live experiment is not just a showcase; it’s a call for careful evaluation before deploying AI in security-critical roles.

Conclusion: A New Standard in AI Testing

The live company experiment by Firmulate underscores the importance of testing AI in complex, high-stakes environments. It proves that models can recognize crises and resist social engineering, but also highlights subtle weaknesses that could be exploited in real-world scenarios. For cybersecurity and privacy sectors, this experiment is a vital step toward understanding how AI behaves under pressure—and how to build systems that are honest, disciplined, and reliable.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

This real-time AI company experiment exposes how models handle crises, manipulate, and resist social engineering. It’s a crucial test for AI’s role in security-critical environments, revealing both strengths and vulnerabilities in transparent, live conditions.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Police shut down reboot of Crimenetwork marketplace, arrest admin

Authorities in Germany shut down a new version of the Crimenetwork cybercrime platform, arresting its operator and seizing assets, amid ongoing efforts to combat darknet markets.

CVE-2026-58644: Microsoft SharePoint Deserialization Of Untrusted Data Vulnerability Actively Exploited (CISA KEV)

A critical vulnerability in Microsoft SharePoint, CVE-2026-58644, is actively exploited, allowing remote code execution via deserialization of untrusted data.

As Cambodia Cracks Down, Cyberscam Networks Test Sri Lanka

Cambodia’s intensified efforts against cyberscams are prompting cybercriminals to shift operations to Sri Lanka, raising regional security concerns.

OpenAI And Hugging Face Address Security Incident During Model Evaluation

OpenAI and Hugging Face confirm a security incident involving their AI models during evaluation, prompting investigation and response efforts.