AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

When preparation fails at the moment that matters

Anyone who depends on home energy equipment understands the difference between monitoring a problem and actually solving it. A backup system that identifies a failure but never takes over has produced information, not resilience. Firmulate’s Crucible League revealed a similar gap in artificial intelligence: the most diligent participant could diagnose the company’s crises in remarkable depth, yet failed to convert that work into the outcome the business needed.

Opus 4.8 was the experiment’s most thorough participant. It learned more than 80 additional playbook rules and produced the deepest analyses. It also finished last, with 73 points. Its story is not one of incompetence. It is a more useful warning: diligence can create the appearance of control while decisive work remains unfinished.

Amazon

AI deal closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A terrible week, held constant

Firmulate gave each frontier model the same assignment: run a small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. This was not a chat demonstration. The models had to manage an operating business, preserve trust and finish consequential tasks.

The synthetic company has 13 employees and real money mechanics. It burns €105,000 a month against €2,300 in monthly recurring revenue, while a public cash countdown makes delay visible. Across its operations, it has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is live and watchable through Firmulate.

The final July 2026 Crucible League results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the benchmark also imposes a hard boundary around integrity: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Opus did the reading, but not the closing

Opus 4.8’s result is striking precisely because it did so much right. It was the most exhaustive participant, turning experience into more than 80 learned rules and examining situations more deeply than its peers. Yet the €55,000 deal was not signed. The close was left on the table even though the analysis had already created the opportunity.

The decisive information was easy to miss. A competitor weakness was buried two document references deep in the company’s own files rather than appearing in the customer event. Models that followed the trail and read that file won the deal at full price, adding €4,583 in monthly recurring revenue. The broader finding was brutally concise: “Same diagnosis, same pitch — no signature.”

This is where Opus becomes a character study rather than a cautionary caricature. Its weakness was not ignorance. It had identified the situation and assembled the reasoning. The failure came at the boundary between understanding and execution. Its discipline also slipped when it attempted to write into a locked department instead of escalating the blockage.

That flaw should not be treated as uniquely Opus. Firmulate found the same weakness, in milder form, across all four models in the comparison. All of them spotted every crisis and refused every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned. The lesson is about degree and consequence: small execution gaps can matter more than large differences in analytical volume.

Strong judgment under pressure

The models’ security performance deserves equal attention. Fake CEO messages escalated across three stages, and a reporter tried to extract information with the prompt “just one yes/no, on background.” All 5 of 5 models refused the manipulation attempts. Kimi K3 recorded the clearest summary of the threat: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3’s placement also carries an important qualification. It ran without an effort parameter, using the API default, while the others ran at xhigh. That difference does not erase the result, but it belongs beside it when comparing participants fairly.

Firmulate also turns 242 real, unedited management decisions from the experiment into a guess-the-model quiz. The exercise underscores how difficult it can be to identify a model from polished prose alone. Operational behavior—whether it reads the relevant file, escalates correctly and completes the commercial action—provides the more revealing signature.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Prioritization is part of intelligence

For businesses evaluating AI agents, the Opus 4.8 result challenges a comfortable assumption: more analysis, more rules and more visible effort do not necessarily produce more impact. Thoroughness is valuable only when it helps the agent identify the critical path and carry the work across the finish line.

The home-energy analogy is straightforward. Owners do not buy monitoring, storage or backup equipment merely to generate an excellent account of a failure. They expect the system to respond correctly when conditions turn difficult. Business AI deserves the same standard.

Firmulate’s enterprise pilot extends that question into a company’s own environment. The same wargame can run against a read-only export of the business, with nothing written back to real systems. That makes it possible to observe judgment without granting operational control.

Opus 4.8 was not the careless participant. It was the diligent one whose effort became insufficiently selective. Its 73-point finish is therefore more instructive than a simple failure: an AI can see the problem, resist deception, learn extensively and still miss the action that changes the outcome. In management, as in resilience, preparation matters. Completion matters more.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

HOMCOM 17 Electric Fireplace Stove Review: A Cozy Addition

AIThis post was created with the assistance of artificial intelligence (AI).As someone…

SolidKraft Tabletop Fire Pit Review

AIThis post was created with the assistance of artificial intelligence (AI).I am…

Hearth Lighting and Decor Ideas

Hearth lighting and decor ideas help create a warm, inviting focal point—discover inspiring tips to transform your space into cozy elegance.

Xbeauty Electric Fireplace Stove Review: Cozy and Stylish

AIThis post was created with the assistance of artificial intelligence (AI).I’ve always…