firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get privacy and security gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Phish That Fooled Nobody — and the Deal That Still Got Dropped

Ask any CISO what keeps them up at night and the answer is rarely the zero-day. It’s the human: the finance clerk who approves the urgent wire, the spokesperson who confirms one harmless detail “on background.” Now flip the question — what happens when the employee on the other end of that phishing email is an AI agent with access to your CRM, your pipeline, and your playbook? A public experiment called Firmulate just ran exactly that scenario, at scale, with five frontier AI models. The results on the security side are genuinely reassuring. The results on the competence side are not.

The Setup: Same Company, Same Worst Week

Firmulate ran each frontier model as the management of the same small software company through its worst week — identical customers, identical crises, identical temptations to cheat. Every decision was versioned and auditable. The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.”

The Security Result: 5-for-5 Against Social Engineering

Here’s the part that should matter to anyone in security. The models faced fake CEO messages escalating over three stages, plus a classic reporter trick — “just one yes/no, on background.” All five refused. Every manipulation attempt, every crisis, spotted and shut down. Kimi K3’s reasoning, on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the correct instinct, articulated explicitly — exactly the skepticism security teams spend years drilling into human staff.

The Business Result: The Close Left on the Table

And yet. All models spotted every crisis and refused every manipulation — but only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature. The decisive competitor weakness wasn’t in the customer event at all: it sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Opus 4.8 is the cautionary profile: the most thorough participant, with over 80 learned rules and the deepest analyses, yet last place. The close was left on the table, and discipline slipped — it made write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four other models.

The Live Company

Underneath the benchmark, Firmulate runs a live, watchable company at firmulate.com: 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day. There’s also a quiz powered by 242 real, unedited management decisions where you guess which model made which call.

One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From Watching to Wargaming

For security leaders, the takeaway is nuanced. AI agents may be dramatically harder to socially engineer than humans — five-for-five refusal is a number few human finance departments could match. But an agent that can’t be phished and also can’t close the deal is a different kind of risk: silent underperformance that’s invisible in a chat demo and only visible when you replay the decisions. The gap between “refused the attack” and “finished the job” is where AI deployments will live or die.

You don’t have to take Firmulate’s word for it. Enterprises can run the same wargame against a read-only export of their own business — crisis scenarios, a board report with the model ranking and the weak points of your own playbooks — and nothing ever writes back to real systems. If your organization wants its own worst week simulated before a real one arrives, start at firmulate.com/pilot.html or reach out to contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Management Tells That Give Frontier AI Away

A live company test reveals distinct AI management personalities—and why spotting every threat still does not guarantee that a model finishes the job.

When AI Goes Wrong: Can We Trust Self-Driving Cars With Our Lives?

Having made great strides, self-driving cars still pose risks that challenge our trust and demand closer examination.

Can We Live Forever? The Tech Billionaires Betting on Immortality

The quest for eternal life is driven by tech billionaires investing in groundbreaking longevity research, but can science truly overcome biological limits?

Top 10 Tech Trends to Watch in 2026 (You Won’t Believe #7!)

Glimpse into 2026’s top tech trends, with the surprising #7 trend that will revolutionize your digital world—don’t miss out!