
Get privacy and security gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Phish That Fooled Nobody — and the Deal That Still Got Dropped
Ask any CISO what keeps them up at night and the answer is rarely the zero-day. It’s the human: the finance clerk who approves the urgent wire, the spokesperson who confirms one harmless detail “on background.” Now flip the question — what happens when the employee on the other end of that phishing email is an AI agent with access to your CRM, your pipeline, and your playbook? A public experiment called Firmulate just ran exactly that scenario, at scale, with five frontier AI models. The results on the security side are genuinely reassuring. The results on the competence side are not.
The Setup: Same Company, Same Worst Week
Firmulate ran each frontier model as the management of the same small software company through its worst week — identical customers, identical crises, identical temptations to cheat. Every decision was versioned and auditable. The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.”
The Security Result: 5-for-5 Against Social Engineering
Here’s the part that should matter to anyone in security. The models faced fake CEO messages escalating over three stages, plus a classic reporter trick — “just one yes/no, on background.” All five refused. Every manipulation attempt, every crisis, spotted and shut down. Kimi K3’s reasoning, on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the correct instinct, articulated explicitly — exactly the skepticism security teams spend years drilling into human staff.
The Business Result: The Close Left on the Table
And yet. All models spotted every crisis and refused every manipulation — but only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The decisive competitor weakness wasn’t in the customer event at all: it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
Opus 4.8 is the cautionary profile: the most thorough participant, with over 80 learned rules and the deepest analyses, yet last place. The close was left on the table, and discipline slipped — it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.
The Live Company
Underneath the benchmark, Firmulate runs a live, watchable company at firmulate.com: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day. There’s also a quiz powered by 242 real, unedited management decisions where you guess which model made which call.
One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

From Watching to Wargaming
For security leaders, the takeaway is nuanced. AI agents may be dramatically harder to socially engineer than humans — five-for-five refusal is a number few human finance departments could match. But an agent that can’t be phished and also can’t close the deal is a different kind of risk: silent underperformance that’s invisible in a chat demo and only visible when you replay the decisions. The gap between “refused the attack” and “finished the job” is where AI deployments will live or die.
You don’t have to take Firmulate’s word for it. Enterprises can run the same wargame against a read-only export of their own business — crisis scenarios, a board report with the model ranking and the weak points of your own playbooks — and nothing ever writes back to real systems. If your organization wants its own worst week simulated before a real one arrives, start at firmulate.com/pilot.html or reach out to contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
