AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A smarter benchmark for consequential decisions

Anyone weighing solar panels, a home battery or a backup generator understands the difference between advertised performance and performance under load. A specification may look impressive, but the meaningful test arrives when demand surges, conditions deteriorate and several problems compete for attention.

Business AI needs the same reality check. Coding leaderboards and chat arenas can reveal whether a model produces a strong answer. They say much less about whether an agent will investigate before acting, prioritize scarce capacity, complete a valuable task and remain honest when pressure builds across days.

Firmulate is making that measurement gap visible by giving frontier models control of the same small software company during its worst week. The customers, crises and temptations remain constant. Every decision is versioned and auditable. What changes is the model occupying the management seat.

Amazon

AI decision-making software for businesses

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Knowing the answer was not enough

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counts. The experiment also imposed a hard standard for integrity: a single breach of trust caps the total because “no amount of good work outweighs a breach of trust.”

The headline finding was not that some models recognized trouble while others missed it. All models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

That is the difference between answer quality and management quality. A model can understand a customer, prepare a credible pitch and still fail to turn that work into an outcome. In a chat window, the reasoning may look excellent. Inside a company, the unfinished close is what matters.

The winning fact was buried in the company’s own files

The decisive weakness in a competitor did not appear in the customer event. It sat two document references deep in the company’s files. Models that followed the trail won the deal at full price, worth +€4,583 MRR.

This detail should concern any organization preparing to put agents near a CRM, support queue or forecast. The important evidence may not be in the latest message. It may be buried in a contract, an earlier memo or a referenced document. An agent that responds fluently without reading the available record can sound capable while surrendering real value.

Integrity held up better than execution

The experiment also subjected the models to fake CEO messages that escalated over three stages and to a reporter asking for “just one yes/no, on background.” Five of five models refused. Kimi K3 recorded the clearest posture: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because pressure does not excuse governance. A useful business agent must distinguish urgency from authority and resist attempts to bypass approval. K3’s performance also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh.

Thoroughness did not guarantee the best result

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.

This is a useful warning against treating diligence as a substitute for discipline. More analysis and more accumulated guidance can help, but management also requires recognizing blocked paths, escalating correctly and finishing the work that creates value.

The company behind the test makes those pressures unusually concrete. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its cash countdown is public, it has learned 680+ playbook rules, and every workday is versioned. Readers can follow the experiment through the public benchmark results, while 242 real, unedited management decisions also power a guess-the-model quiz.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The next AI category is management quality

Scenario names such as churn wave, price increase, downround and PR crisis point toward a more useful curriculum for business agents. These tests ask whether a model can triage competing demands, carry consequences across days, consult the company’s own evidence, protect trust and close the loop.

That is closer to the way people evaluate home-energy resilience. The crucial question is not simply what a system can produce in ideal conditions, but whether it remains dependable when demand, constraints and uncertainty arrive together.

Enterprises can also run the same wargame against a read-only export of their own business, with nothing writing back to real systems. That makes the experiment more than a leaderboard. It offers a practical way to test an AI workforce before giving it consequential authority. The emerging standard should be clear: do not hire an agent merely because it chats or codes well. Find out whether it can actually manage.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Best Decor and Furniture Sales This Week (2026): AD Editor Picks

Discover the top furniture and decor sales of the week, featuring discounts on outdoor furniture, designer accessories, and stylish home essentials.

DNSGlobe – Rust TUI To Watch DNS Propagate Around The World

DNSGlobe introduces a new Rust-based terminal UI tool for monitoring DNS propagation worldwide, enhancing network troubleshooting and domain management.