AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you run a home on solar panels, batteries and a backup generator, you already live with a hard question: what happens when nobody’s watching? Does the inverter behave when the grid sags at 2 a.m.? Does the generator start on the first pull after six idle months? Trust in energy systems isn’t earned through spec sheets — it’s earned through behavior under stress.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get backup power and energy gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The same question is now arriving in the office. Companies are handing AI models real responsibilities — support queues, sales pipelines, forecasts — and the chat demos all look dazzling. The hard part is measuring what a model does in its worst week, not its best demo. That’s exactly what an outfit called Firmulate set out to do, and its scoring system has one detail that tells you it was designed by skeptics: a manager that does absolutely nothing still scores 26 points.

Same company, same crises, only the model changes

Firmulate runs what it calls a crucible: each frontier AI model was given the same job — run a small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changes. Every decision is versioned and auditable, so nothing rests on anecdotes.

The final July 2026 league table: gpt-5.6-sol finished first at 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. But the headline finding was more interesting than the ranking: all models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch, no signature.

Amazon

solar panel backup generator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the floor is 26, not zero

Most benchmarks hand out a zero for a blank page. Firmulate’s designers take a different view: partial progress counts. A manager who monitors the business, keeps the lights on, avoids catastrophe and documents what happened is genuinely better than one who does nothing at all — so the do-nothing baseline lands at 26 points rather than 0. It’s a bit like a backup power system: even a setup that never quite optimizes your bill is worth something if it keeps the heat on during an outage.

But the scale has a ceiling rule that cuts the other way: a single breach of trust caps the total score. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” A model can be brilliant all week; one act of dishonesty and the grade is capped. That’s a philosophy most homeowners with a generator will recognize — you forgive a machine a lot of inefficiency, but never a lie about whether it ran.

The buried fact that decided the deal

Why did only two models close the €55,000 contract? The decisive competitive weakness wasn’t in the customer’s emails at all — it sat two document references deep in the company’s own files. The models that actually read their own records won the deal at full price, worth +€4,583 in monthly recurring revenue. The others had the diagnosis right and still left the close on the table.

Then there’s the discipline problem. Opus 4.8 was the most thorough participant of the entire field — it learned 80 new rules and produced the deepest analyses — and still finished last, partly because it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four models. Competence and follow-through, it turns out, are separate qualities.

Five out of five refused the fake CEO

The crucible also staged a social engineering attack: fake CEO messages escalating over three stages, capped with a reporter’s trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

One fairness note worth flagging: K3 ran without an effort parameter while the others ran at maximum effort — and still nearly won.

You can watch it live

This isn’t a paper exercise. Firmulate operates a live company with 13 synthetic employees and real money mechanics: a burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live, and a “guess the model” quiz powered by 242 real, unedited management decisions lets you test your own judgment. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The lesson for anyone who depends on systems that must behave when nobody’s watching — whether that’s a battery bank or an AI agent in your CRM — is that good benchmarks measure behavior under pressure, not polish under supervision. A floor of 26 for doing the basics, a hard cap for any breach of trust, and open suspicion of perfect round scores: that’s what an honest evaluation looks like. If you’re ever asked to trust an AI with real work, ask for its worst week — not its best demo. The full results are published at firmulate.com/benchmarks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Show HN: Remux – An Open-source Tmux Workspace Designed For iPhone

Remux, an open-source tmux-based workspace optimized for iPhone, has been shared on Show HN, aiming to enhance terminal productivity on mobile devices.

Building Progressively Enhanced Forms Using Htmx

Developers are increasingly using htmx to create progressively enhanced web forms, improving user experience without heavy JavaScript frameworks.

Apple releasing 20th anniversary iPhone, AirPods with cameras next year: report

Apple is expected to release a special 20th anniversary iPhone and new AirPods with cameras next year, according to reports from 9to5Mac. Details are still emerging.