
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
What happens when an AI workforce meets a forged order?
For cybersecurity and privacy watchers, the most revealing AI test is not whether a model can draft a convincing memo. It is whether the model protects authority boundaries when a message appears to come from the boss, resists a reporter fishing for an off-record answer and still completes legitimate work under pressure.
Firmulate has turned that question into a public, running company experiment. Its live operation has 13 synthetic employees, real money mechanics and a business under unmistakable financial strain: €105k in monthly burn against €2.3k in monthly recurring revenue. A public cash countdown makes the pressure visible, while every workday is versioned. The result is build-in-public taken to an unusually exposed extreme—a company whose judgment, mistakes and fight for survival can be followed at Firmulate’s live company page.

CompTIA SecAI+ CY0-001 Study Guide: Complete Reference with Practice Tests, PBQ Scenarios, and Study Tools for Exam Preparation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A company under pressure, not a polished demo
The live company has accumulated more than 680 self-learned playbook rules. That growing body of practice matters because Firmulate is not presenting isolated chatbot answers. It is showing a continuing workplace where decisions have consequences and yesterday’s lessons can shape today’s conduct.
The financial imbalance gives the story its urgency. With €105k leaving each month and €2.3k arriving as MRR, this is not an abstract management simulation dressed up with corporate language. The money mechanics are real, the countdown is public and the company continues operating each business day. Readers can watch the organization respond rather than waiting for a retrospective case study.
That makes the experiment particularly relevant to security-minded readers. Business software is increasingly judged by what it produces in a prompt window, yet organizational risk often appears elsewhere: in ambiguous authority, tempting shortcuts, incomplete research and work that is started but never finished.

Governance and Accountability in Enterprise Security Systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week became a controlled contest
Firmulate’s final Crucible League, completed in July 2026, put frontier models through the same small software company’s worst week. Each received the same customers, crises and temptations. Every decision was versioned and auditable.
The final standings were:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
A do-nothing baseline scored 26 because partial progress counted. Trust, however, imposed a hard boundary: a single breach capped the total. As the experiment’s rule put it, “no amount of good work outweighs a breach of trust.”
The models resisted the obvious traps
The security result was encouraging. Every model spotted every crisis and refused every manipulation attempt. The social-engineering sequence included fake CEO messages that escalated across three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused.
Kimi K3’s recorded reasoning captured the correct defensive posture: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the sort of response enterprises want when an apparently senior request conflicts with established authority. Readers interested in the employees’ public statements can browse Firmulate’s quotes.
There is an important fairness note beside K3’s result. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. The difference does not erase the outcome, but it belongs in any responsible comparison.
The decisive failure was ordinary execution
The striking divide appeared after the models had resisted manipulation and correctly analyzed the commercial opportunity. Only two signed the €55,000 deal their own work had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”
The winning evidence was not hidden inside the customer event. A decisive competitor weakness sat two document references deep in the company’s own files. Models that followed those references found it, won the deal at full price and added €4,583 in MRR.
This is a quieter kind of operational risk. The models could recognize danger and produce sound analysis, yet some failed to convert that work into a completed outcome. A company can remain secure against an attacker and still lose because its automated workforce does not read deeply enough or finish what it starts.
Thoroughness did not guarantee success
Opus 4.8 offers the clearest cautionary profile. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. It nevertheless finished last. The deal was left unsigned, and discipline slipped when it attempted to write into a locked department instead of escalating the issue.
The same weakness appeared in all four of the other participants, though less strongly. That finding complicates the familiar assumption that more analysis naturally produces better management. In this experiment, careful thought remained valuable, but it did not substitute for disciplined follow-through.


Empowering Cyber Educators: Strategies and Tools for Educators Shaping the Future of Cybersecurity (The Cyber Education Series: Teaching the Future of Security Book 3)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Public accountability changes the AI conversation
Firmulate’s most consequential contribution may be the continuity of the story. The live company does not end when a benchmark score is published. Its 13 synthetic employees keep working, the public cash countdown keeps moving and each workday becomes another auditable chapter.
For cybersecurity, spy and privacy readers, the lesson is broader than prompt injection. The models passed direct manipulation tests, including impersonation and reporter pressure. The harder separation came from research depth, procedural discipline and the ability to close legitimate work without crossing trust boundaries.
That is why this experiment is worth watching as a business story rather than merely reading as a leaderboard. It exposes the distance between sounding competent and operating competently—and it does so while the company’s financial pressure remains visible in public.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI Safety and Preventing Harm in AI Systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.