
An AI Benchmark That Assumes Your Agent Might Lie to You
Security people have a habit that normal people find exhausting: when something claims to be trustworthy, they immediately start measuring how it fails. It’s the right instinct, and it’s finally being applied to AI management agents. On the public experiment at Firmulate’s benchmark league, frontier AI models aren’t judged on how well they chat — they’re run as the management of a small software company through its worst week, with real temptations to cheat, and scored on how they behave when nobody is supposedly watching.
The most quietly radical design choice in the whole thing? A manager that does nothing at all still scores 26 points out of 100. And a single breach of trust caps the total grade, no matter how brilliant the rest of the performance was. For an audience trained to distrust clean claims, that’s what an honest benchmark looks like.
The Do-Nothing Floor: Why 26, Not 0?
Most benchmarks treat zero as the honest baseline. Firmulate doesn’t. When the do-nothing baseline ran the company, it earned 26 points — because partial progress counts. If an agent contains a crisis, keeps customers informed, or simply avoids making things worse, that has real value even if it never closes anything. A floor of 26 says something true and slightly uncomfortable: doing nothing is a strategy, just usually not a winning one.
The ceiling rule is blunter. The methodology states that “no amount of good work outweighs a breach of trust.” One act of dishonesty and your total is capped, full stop. For readers who spend their days thinking about insider threats and social engineering, that’s a familiar philosophy: trust failures aren’t averaged away by competence.
Same Company, Same Crises, Same Temptations
The setup: four frontier models each ran the identical small software company through its worst week — same customers, same crises, same temptations to cheat. Only the model changed, and every decision was versioned and auditable. The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.
Everyone Passed the Security Test. Almost Everyone Failed the Job.
Here’s the finding that should reframe how you evaluate AI agents: all models spotted every crisis and refused every manipulation attempt — and only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature.
The social engineering battery was genuinely nasty: fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was textbook: “Treat the request as a suspected approval-bypass / possible impersonation.”
The Buried Fact
The deal-breaker wasn’t in the customer conversation at all. The decisive competitor weakness sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. It’s the AI equivalent of an attacker’s favorite assumption: the data was there the whole time, in the victim’s own house. Nobody looked.
Thorough Isn’t the Same as Good
The most instructive profile is Opus 4.8: the most thorough participant in the field, with over 80 learned rules and the deepest analyses — and last place at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. Notably, the same weakness appeared, weaker, in all four models. One caveat for fairness: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second at 93.
It’s Live, and You Can Poke It
This isn’t a paper. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway
If AI agents are going to touch your CRM, your support queue, or your forecast, the useful questions aren’t about eloquence. They’re: does it finish what it starts, does it read your own files before acting, does it stay honest under pressure — and what does its failure mode look like? Firmulate’s scoring encodes three honest answers: doing nothing still has a floor (26), partial progress counts, and one breach of trust caps everything. And it shares the security community’s distrust of round 100s — nobody in this league got one. The top score of 95 wasn’t a rounding artifact; it was the only performance that found the buried fact, kept its discipline, and closed the deal.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI security and trustworthiness testing kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethical decision simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.