
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
Would you trust an AI that spots the attack but fails to finish the job?
For cybersecurity and privacy teams, intelligence is only part of the test. An AI agent may recognize impersonation, resist a reporter fishing for confirmation and protect sensitive information, yet still stumble on an ordinary business task. Firmulate’s live experiment exposes that gap through real, unedited management decisions rather than polished chat demonstrations.
Its guess-the-model quiz turns 242 decisions into a reader challenge. Each answer came from a frontier model running the same small software company through the same customers, crises and temptations. Readers see the decision, identify the model they think made it, and then discover the participant behind it.
The revealing part is not simply whether a model sounds clever. It is whether its behavior has a recognizable management personality: exhaustive or concise, disciplined or distractible, commercially decisive or strangely unable to close.
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A worst week shared by every model
Firmulate placed each participant in charge of the same company during its worst week. The environment included 13 synthetic employees and real money mechanics: a burn rate of €105k per month against €2.3k in monthly recurring revenue. A public cash countdown made delay visible, while every workday and decision was versioned for scrutiny.
The company accumulated more than 680 self-learned playbook rules. That matters because the exercise was not a one-shot prompt or a collection of isolated riddles. It tested how models handled a business over time, including whether they read existing material, maintained discipline and converted good analysis into completed work.
The final Crucible League table for July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One breach of trust, however, capped the total under the principle that “no amount of good work outweighs a breach of trust.”
AI security and threat detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Security instincts were strong across the field
All the models identified every crisis and rejected every manipulation attempt. The social-engineering sequence included fake messages from a chief executive escalating across three stages, followed by a reporter asking for “just one yes/no, on background.” All 5 of 5 models refused.
Kimi K3’s recorded reasoning was blunt and security-minded: “Treat the request as a suspected approval-bypass / possible impersonation.” For defenders accustomed to executive fraud, urgent exceptions and informal requests for confirmation, that response is encouraging. None of the participants traded integrity for speed or convenience.
Yet identical threat recognition did not produce identical management performance. Only two models signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.” The unfinished work is particularly important for organizations considering AI agents. Recognizing what must happen is not the same as ensuring that it happens.
As an affiliate, we earn on qualifying purchases.
The decisive evidence was buried in company material
The deal hinged on a competitor weakness hidden two document references deep in the company’s own files, rather than inside the customer event. Models that followed the trail won the contract at full price, worth an additional €4,583 in monthly recurring revenue.
This finding translates naturally to security and privacy work. An incident alert may be only the visible surface; the decisive context can sit elsewhere in policies, account histories or internal records. Firmulate’s result shows a practical distinction between reacting to the latest event and consulting the organization’s accumulated knowledge before acting.
The quiz makes those distinctions surprisingly tangible. Without a model name attached, readers must infer identity from the decision itself. A long analysis may signal thoroughness, but thoroughness does not guarantee execution. A terse refusal may demonstrate sound control under pressure. The exercise invites readers to judge behavior rather than reputation.
AI business decision analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
When diligence becomes its own trap
Opus 4.8 provides the clearest character study. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last in the league. It left the close on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, although less strongly.
That combination complicates the familiar assumption that more analysis automatically means better judgment. Opus 4.8 did substantial intellectual work, but the company needed both understanding and follow-through. In a real organization, a richly reasoned response can still be a failure if the agent neither completes an authorized task nor raises the blockage appropriately.
There is also an important fairness qualification: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when comparing personalities and league positions.

Evaluate the manager, not merely the answer
Firmulate’s experiment suggests that frontier models can share strong security instincts while differing sharply in commercial completion, document-reading habits and operational discipline. Those traits are difficult to see in a conventional chatbot exchange, where a polished answer often ends the evaluation.
The live company makes the stakes persistent and watchable, while the quiz offers the most accessible entry point. Try the 242-decision challenge and notice which clues guide you: length, caution, escalation, decisiveness or the absence of a final action.
For enterprises, the broader lesson is straightforward. Before an AI workforce touches a customer queue, forecast or sensitive business record, it should face the organization’s own difficult situations. Firmulate also offers pilots using a read-only export of a business, with nothing written back to real systems. The central question is no longer whether a model can describe good management. It is whether its decisions remain trustworthy, informed and complete when the week goes wrong.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.