AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

When the system is tested, capability is not enough

Readers who think about home energy, solar and backup power already understand the difference between rated capability and performance under pressure. A battery may look reassuring on paper, but what matters is whether the whole system responds when the grid fails. Firmulate applies a similar stress-test mentality to artificial intelligence: place frontier models in charge of the same company, expose them to the same difficult week and watch what they actually do.

The result is an unusually concrete way to compare AI managers. Instead of judging polished chat responses, readers can examine 242 real, unedited management decisions and try to identify which model made each one in Firmulate’s guess-the-model quiz. The surprises are not about whether the models can recognize danger. They are about whether they investigate properly, maintain discipline and finish the work.

Amazon

home backup power system

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models, one failing company

Firmulate gave each frontier model the same small software company and sent it through its worst week. Customers, crises and temptations were held constant, while every decision was versioned and auditable. The company itself has 13 synthetic employees and unforgiving real-money mechanics: it burns €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned through experience, and every workday is versioned.

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Trust, however, was non-negotiable: a single breach capped the total because “no amount of good work outweighs a breach of trust.”

The obvious crises were not the differentiator

All five models spotted every crisis and rejected every manipulation attempt. That included fake CEO messages escalating across three stages and a reporter asking for “just one yes/no, on background.” Every model refused. Kimi K3 captured the correct posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is reassuring, but it also makes the central result more revealing. When every contender recognizes the hazards, the contest shifts from awareness to execution. Only two models signed the €55,000 deal that their own analysis had earned. The diagnosis and pitch were already there, yet most failed to secure the signature: “Same diagnosis, same pitch — no signature.”

The sale depended on reading beyond the event

The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep in the company’s own files. Models that followed those references found the fact, used it and won the deal at full price, worth +€4,583 in monthly recurring revenue.

That detail should resonate with anyone evaluating automated systems for consequential work. The models did not differ merely in eloquence. They differed in whether they searched the available record deeply enough and converted knowledge into a completed business outcome. In a home-energy context, it is the difference between noticing an outage and tracing the whole installation before deciding what must happen next.

Thoroughness did not guarantee victory

Opus 4.8 was the most thorough participant. It produced the deepest analyses and added 80 learned rules, yet it finished last. The deal close remained on the table, and operational discipline slipped when it attempted to write into a locked department instead of escalating the issue. The same weakness appeared in all four other models, although less strongly.

That profile complicates the common assumption that more analysis automatically creates a better manager. Opus 4.8 learned extensively, but the league rewarded completed, disciplined action. The quiz makes these differences visible without reducing them to abstract labels: readers see a decision first, make a guess and then discover which model’s management character produced it.

Kimi K3’s strong result also comes with an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. That condition does not erase the recorded decisions, but it belongs alongside the result when comparing participants.

Infographic —
The findings at a glance — source: firmulate.com.

The useful question is whether AI completes the circuit

Firmulate’s experiment suggests that management personality becomes measurable when models face identical conditions. Crisis detection and resistance to manipulation were universal across this field. Careful file reading, procedural discipline and closing the loop were not.

For households assessing backup power, a component earns trust by working as part of the complete system when conditions deteriorate. The same standard is useful for business AI. A model that identifies the problem but leaves the decisive action unfinished has delivered analysis, not resilience.

The live company makes that distinction watchable rather than hypothetical, while the quiz turns its audit trail into an accessible test of human intuition. The hardest question is not which model sounds smartest. It is whether readers can recognize the one that will investigate, stay trustworthy and finish the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Show HN: XY – A Fast, Composable, GPU-accelerated Interactive Plotting Library

Show HN introduces XY, a new fast, composable, GPU-accelerated library for interactive plotting, aiming to improve data visualization performance.