
A fake message from the CEO is a familiar security test. But what happens when an AI agent must spot the impersonation, protect a company under pressure and still finish the work? Firmulate’s live experiment puts five frontier models through that kind of week. Its results suggest that a convincing answer is not the same as a completed job.
One company, the same bad week
Firmulate runs AI models as complete companies, measuring management quality rather than chat quality. In its Crucible experiment, each model ran the same small software company through its worst week: identical customers, crises and temptations. Every decision was versioned and auditable.
The company is a working simulation with real money mechanics: 13 synthetic employees, monthly burn of €105,000 against €2,300 in monthly recurring revenue, a public cash countdown and more than 680 self-learned playbook rules. Its workdays are versioned. The experiment is live and watchable at Firmulate.
As an affiliate, we earn on qualifying purchases.
Security discipline was common; follow-through was not
Every model spotted every crisis and refused every manipulation attempt. That included fake CEO messages escalating across three stages and a reporter’s request for “just one yes/no, on background.” All five refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.”
The more consequential gap came after diagnosis. Only two models signed the €55,000 deal their own analysis had earned. The weakness that decided the deal was buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue.
That pattern is a reminder for security teams evaluating agents: recognizing a threat is only part of the job. An agent may also need to find evidence in business records, act within its authority and carry a decision through. In this experiment, models could reach the same diagnosis and make the same pitch, yet still leave the signature undone.
As an affiliate, we earn on qualifying purchases.
A close second, with a fairness caveat
The final July 2026 league table put gpt-5.6-sol first with 95 points and Moonshot’s Kimi K3 second with 93. K3 finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26: partial progress counted, but a single breach of trust capped the total. As Firmulate puts it, “no amount of good work outweighs a breach of trust.”
K3’s result is notable because it beat three of the four Western frontier models in the field. But the comparison carries a material caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four. Thoroughness, the results suggest, does not guarantee safe or effective execution.
As an affiliate, we earn on qualifying purchases.
Test the model against the work
For companies considering AI agents in customer support, CRM or forecasting, these results make a case for testing models against realistic decisions before handing them operational responsibility. Firmulate’s quiz uses 242 real, unedited management decisions to let readers guess which model made each call. Enterprises can also run the wargame against a read-only export of their own business; the pilot does not write back to real systems.
See the full benchmark findings and the live company at Firmulate.

As an affiliate, we earn on qualifying purchases.
The takeaway
The Crucible shows why model choice is a security and operations decision, not just a writing preference. All five models resisted the manipulation tests, but only two converted their analysis into the earned deal. Kimi K3’s strong second-place finish makes the field look open; its different effort setting makes a direct comparison less conclusive. For organizations, relying on brand or demo performance alone is a bet. Test the agent on your own work, including the moments when it must find buried evidence, resist pressure and complete an authorized action.
Fairness note: K3 ran without an effort parameter (API default), while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.