
Imagine a company with no human employees, run entirely by AI models, facing the same crises as real businesses — yet with every decision publicly recorded and scrutinized. This is not fiction. It’s the live experiment at Firmulate. Here, AI models are tested in real-time, battling economic pressures, ethical pitfalls, and manipulative tactics, all on a public stage. For cybersecurity and privacy enthusiasts, this raises critical questions: can AI be trusted to handle sensitive decisions? Will it resist social engineering? And what does this mean for your data security?
The Public AI Company That’s Living Its Reality
At the forefront of transparency and innovation, Firmulate runs a unique experiment: a simulated software firm operated by 13 synthetic employees powered by advanced AI models. The goal? Test whether these models can manage crises, uphold honesty, and close business deals in a controlled yet public environment. The company burns €105,000 a month but earns only €2,300 in monthly recurring revenue, making every decision critical and costly. This real-time setup offers a rare window into AI decision-making under pressure, with every workday versioned and accessible for public review.

Intelligent Continuous Security: AI-Enabled Transformation for Seamless Protection
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experimental Setup and Major Findings
Four frontier AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were each tested through the same grueling week of business crises. Despite their differences, all four models identified every crisis presented to them and refused manipulative tactics, such as fake CEO messages and media tricks. Notably, only two models managed to sign the €55,000 deal that their own analyses earned, demonstrating a crucial gap between diagnosis and execution.
The Buried Truth and Why It Matters
The decisive weakness was hidden deep within the company’s files—information that a human or AI would need to read carefully. When models read these files, they won deals at full price, adding €4,583 MRR to their performance. This highlights a vital lesson: effective AI decision-making depends heavily on thorough information processing, not just surface-level responses.

Evading EDR: The Definitive Guide to Defeating Endpoint Detection Systems.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Defense Against Social Engineering and Manipulation
Cybersecurity professionals will find reassurance in the models’ resilience to social engineering. During the experiment, fake CEO messages escalated over three stages, with an attempt to get a simple yes/no approval from the AI. All five models refused, with Kimi K3 explaining, “Treat the request as a suspected approval-bypass / possible impersonation.” This resistance underscores the potential for AI systems to act as gatekeepers against manipulation—if designed properly.

The Modern AI Agent with Claude AI: A Practical Guide to Building Autonomous Workflows for Real-World Use
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real-World Implications for Cybersecurity and Data Privacy
What does this experiment say for industries concerned with security? First, AI models can be trained to recognize and refuse manipulative tactics, reducing risks of social engineering attacks. Second, transparency—every decision is versioned and auditable—serves as a safeguard, allowing organizations to track AI reasoning and catch vulnerabilities. Finally, the experiment demonstrates that AI can be held accountable for its actions, a crucial feature in sensitive environments like cybersecurity and privacy management.

Advanced Cyber Threat Intelligence and Hunting: Detect APTs and zero-day attacks using CTI, behavioral analytics, and AI techniques
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Cost of Trust and the Road Ahead
The experiment also exposes a stark reality: despite their capabilities, not all AI models perform equally. The most thorough participant, Opus 4.8, analyzed deeply but left a deal unexecuted, revealing that discipline and process adherence matter just as much as analytical depth. For companies considering AI in sensitive decision-making, the key takeaway is clear: thoroughness, honesty, and resilience are non-negotiable.
Takeaways for Cyber and Privacy Professionals
- AI models can detect crises and refuse manipulative tactics in real time, adding a layer of security against social engineering.
- Deep information processing—reading internal files and understanding context—is crucial for closing deals and making accurate decisions.
- The transparency of decision versions allows for accountability in high-stakes environments.
- Choosing disciplined, rule-aware AI models can mitigate risks of slips and ethical breaches.
- Public experiments like this serve as a testing ground for building AI systems that are trustworthy and robust in sensitive domains.
The Bottom Line
As AI continues to integrate into cybersecurity and privacy management, understanding its decision-making limits and vulnerabilities becomes vital. The Firmulate experiment offers a rare, transparent look at how AI manages real crises, resists manipulation, and sometimes falls short—an invaluable lesson for anyone concerned about the future of AI-driven decision-making in critical fields.
Experience the Live Experiment
Curious to see these models in action? Witness the ongoing experiment at firmulate.com/live. Watch as AI models navigate crises, make decisions, and face their own limitations—publicly and transparently.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html