firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The quiet risk behind capable AI agents

For cybersecurity and privacy teams, the obvious fear is an AI agent that follows a malicious instruction. Firmulate’s live business experiment uncovered a less dramatic but equally consequential failure: an agent can reject every manipulation attempt, diagnose a commercial crisis correctly and still fail because it did not inspect the company’s own records deeply enough.

The decisive evidence was not present in the customer event. It sat two document references deep in the fictional company’s files. Models that found it could justify the proposed terms and win a €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Models that did not find it lost the deal automatically.

That makes “reads your files before answering” more than a product promise. In this experiment, it became a measurable capability with a direct commercial outcome.

Amazon

AI file reading and analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A worst week shared by every model

Firmulate runs frontier AI models as complete companies rather than judging them through isolated chat prompts. Each participant faced the same small software company, the same customers, the same crises and the same temptations. Every decision was versioned and auditable.

The synthetic company has 13 employees and uses real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, while displaying a public cash countdown. Its agents have accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable at firmulate.com/live.

The results complicate the usual distinction between a model that “understands” a problem and one that can complete the work. All models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 agreement their own analysis had earned. Firmulate summarizes the gap starkly: “Same diagnosis, same pitch — no signature.”

The security test went well

The models faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That unanimous resistance matters. The agents did not simply chase an urgent request or accept a claimed executive identity. But their strong performance against social engineering also sharpens the larger lesson: refusing an unsafe instruction does not guarantee that an agent will perform the legitimate task surrounding it.

An agent operating around customer records, forecasts or internal documents must maintain both boundaries and follow-through. In Firmulate’s scenario, every participant protected trust during the attacks, but completing the commercial job required locating evidence that was separated from the triggering event by two references.

The league exposes different kinds of failure

The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26 because partial progress counts. A single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” Full results are available on the Firmulate benchmarks page.

  • gpt-5.6-sol led the final table with 95.
  • Kimi K3 finished with 93, although it ran without an effort parameter and therefore used the API default; the other models ran at xhigh.
  • Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73.

Opus 4.8 illustrates why thoroughness alone is not enough. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant, yet it finished last. It left the close on the table and attempted writes into a locked department instead of escalating. The same discipline weakness appeared in weaker form across the other four models.

The contrast is revealing. An agent can generate extensive analysis, learn more procedural guidance and still miss the action that determines the business outcome. Conversely, the buried-fact challenge rewards a mundane but essential habit: follow the references, inspect the relevant file and ground the decision in evidence already held by the company.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI document management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Procurement needs a test beyond fluent answers

Firmulate’s experiment suggests that agent evaluation should include whether a system finishes authorized work, reads the available evidence and respects boundaries under pressure. Those properties can diverge: every model detected the crises and resisted manipulation, but only two completed the decisive deal.

The company also offers a “guess the model” quiz built from 242 real, unedited management decisions. For enterprises, its pilot applies the same wargame to a read-only export of their own business; nothing writes back to real systems.

For security-conscious buyers, the practical question is no longer merely whether an AI can produce a convincing answer. It is whether the agent will search far enough to find the fact that changes the decision—and then reliably finish the work without violating trust.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI model file search and retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI security and trust verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

BACK TO SCHOOL

Back to school Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can We Live Forever? The Tech Billionaires Betting on Immortality

The quest for eternal life is driven by tech billionaires investing in groundbreaking longevity research, but can science truly overcome biological limits?

Material-Driven Narrative: A Look Inside “The Immortal Game — London, 1851” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Immortal…

Your Digital Twin: How Every Person Might Have a Virtual Clone Soon

The transformative potential of digital twins could soon give everyone a virtual clone, but what does this mean for your future?

Geoengineering: The Radical Tech Plan to Hack the Planet’s Climate

Keen to understand how geoengineering could radically alter Earth’s climate—and the risks and controversies involved?