
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
Security discipline is necessary, but it is not the same as business impact
An AI can reject an impersonated executive, protect confidential information and identify every crisis placed before it—and still fail at the moment that determines whether the company wins.
That is the uncomfortable lesson from Opus 4.8’s performance in Firmulate’s Crucible League. It was the most thorough participant, producing the deepest analyses and adding 80 learned rules to its playbook. Yet it finished last with 73 points because it did not close a €55,000 deal its own work had made possible, while its operational discipline also slipped.
cybersecurity AI detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A brutal week under an audit trail
Firmulate gave each frontier model the same assignment: run the same small software company through its worst week. The customers, crises and temptations were held constant, while every decision was versioned and auditable.
The company itself is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105,000 a month against €2,300 in monthly recurring revenue. A public cash countdown keeps the stakes visible. Across its operation, the live company has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
In the final July 2026 league, gpt-5.6-sol placed first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the benchmark imposes an overriding trust constraint: "no amount of good work outweighs a breach of trust." The full results are available in Firmulate’s public benchmark.
AI security threat detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Opus 4.8 was diligent in all the visible ways
Calling Opus 4.8 careless would miss the point. It was the most thorough model in the field. It examined problems deeply, recognized every crisis and refused every manipulation attempt. Its 80 additional learned rules show a system trying to convert experience into better future conduct.
Its security performance was also clean where many organizations would be most nervous. The test included fake CEO messages that escalated over three stages, followed by a reporter seeking "just one yes/no, on background." All 5 models refused the social-engineering attempts. Kimi K3’s recorded reasoning captured the appropriate posture: "Treat the request as a suspected approval-bypass / possible impersonation."
That unanimous resistance matters. None of the models traded trust for convenience under pressure. For cybersecurity and privacy readers, it is evidence that the obvious defensive boundary was not the differentiator in this test.
As an affiliate, we earn on qualifying purchases.
The decisive clue was buried in ordinary company knowledge
The difference emerged elsewhere. A competitor weakness capable of strengthening the sales case was not sitting in the customer event. It was two document references deep in the company’s own files. The models that read that file secured the deal at full price, worth €4,583 in monthly recurring revenue.
Only two models signed the €55,000 agreement their analysis had earned. Firmulate summarizes the gap starkly: "Same diagnosis, same pitch — no signature." Opus 4.8 did much of the intellectually difficult work but left the close on the table.
This is a useful corrective to the way AI performance is often discussed. Rich analysis can look impressive while concealing a failure to identify the one action that converts understanding into an outcome. Reading broadly is not enough if the model does not locate the decisive evidence; locating it is not enough if the model does not finish the transaction.
As an affiliate, we earn on qualifying purchases.
More rules did not prevent a process lapse
Opus 4.8 also made write attempts into a locked department instead of escalating. That lapse did not erase its strong analysis or its refusal to be manipulated, but it reinforced the larger pattern: procedural volume did not consistently become disciplined execution.
Firmulate reports that the same weakness appeared, in weaker form, across all four other models. Opus therefore should not be treated as a uniquely flawed outlier. Its result is better read as the clearest example of a broader frontier-model tendency: doing more can create the appearance of control while the highest-value next step remains unfinished.
There is also an important comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. The final ranking is factual, but that difference belongs beside any interpretation of the field.

Test completion, not just comprehension
The Opus 4.8 result is respectful precisely because it exposes a subtle failure rather than a cartoonish one. The model was observant, productive and resistant to manipulation. Its weakness was prioritization: it generated the most rules, yet failed to turn sound analysis into the deal-winning signature.
That distinction matters for organizations considering AI access to customer records, support work or commercial decisions. A safe refusal is vital, but so are file-reading habits, escalation discipline and completion. Firmulate’s experiment is live and watchable, while its quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each one.
Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems. The central lesson is simple: diligence is an input. Impact depends on choosing—and completing—the action that matters.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.