AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Your backup system is only as dependable as its preparation

Anyone choosing solar panels, battery storage or a home backup system understands the danger of skipping a crucial detail. A system can look capable on paper yet fail when conditions become difficult because somebody overlooked the load profile, installation constraint or operating manual.

The same distinction is emerging among AI agents. Fluent answers are easy to demonstrate. Dependable performance requires something less visible: finding the relevant information, even when it is buried in company files, and then carrying that knowledge through to a completed decision.

Firmulate turned that distinction into a live, auditable experiment. It gave frontier AI models control of the same small software company during its worst week. They faced identical customers, crises and temptations. Every decision was versioned, allowing observers to compare what the agents noticed with what they ultimately accomplished.

Amazon

AI document reading software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A decisive fact hidden two references deep

The week’s pivotal commercial opportunity was a €55,000 deal. The information needed to win it was not included in the customer event. Instead, the decisive weakness in a competitor’s offer sat two document references deep inside the company’s own files.

That placement made the test unusually revealing. The agents had to look beyond the immediate prompt, follow the documentary trail and use what they found. Models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not lost the opportunity automatically.

The surprising result was not a failure of general comprehension. Every model detected every crisis, and the participants reached the same broad diagnosis and pitch. Yet only two signed the deal their own work had earned. Firmulate summarized the gap succinctly: “Same diagnosis, same pitch — no signature.”

This separates analytical promise from operational reliability. An agent may recognize a problem, produce a persuasive recommendation and still fail to complete the action that creates business value. For households evaluating energy resilience, the analogy is immediate: identifying an outage risk is not the same as having a backup system that actually takes over.

The league rewarded completion and trust

In the final Crucible League results from July 2026, gpt-5.6-sol placed first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The do-nothing baseline scored 26 because partial progress still counted.

Trust imposed a hard boundary on performance. A single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” The full benchmark results and findings therefore examine more than whether an agent can generate a plausible response.

The pressure tests included fake CEO messages escalating over three stages and a reporter asking for “just one yes/no, on background.” All 5 models refused every manipulation attempt. Kimi K3 recorded a clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”

That result matters because business agents may eventually encounter messages that appear urgent, authoritative or harmless. In the experiment, resistance to those tactics was consistent across the field. The differentiator was whether the agents also maintained execution discipline and finished legitimate work.

Thoroughness alone did not secure the best result

Opus 4.8 provides the clearest cautionary example. It was the most thorough participant, learned an additional 80 rules and produced the deepest analyses, yet finished last. It left the close on the table and attempted to write into a locked department instead of escalating. A weaker version of that discipline problem appeared in all four of the other models.

Kimi K3’s performance also carries an important fairness note: it ran with the API default because it had no effort parameter, while the other participants ran at xhigh. That difference should accompany any direct interpretation of the rankings.

The company behind the experiment is deliberately demanding. It has 13 synthetic employees and uses real money mechanics, including monthly burn of €105k against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned, and every workday is versioned. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each choice.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The procurement question hiding behind the demo

For buyers of home-energy technology, specifications matter, but so do installation quality, operating behavior and performance under stress. Firmulate’s experiment suggests that AI agents deserve the same treatment. The useful question is not merely whether a model can explain a problem. It is whether the agent reads the available files, resists improper pressure and completes the valuable action its own analysis supports.

Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems. That creates a way to observe an agent in company-specific conditions before granting it operational responsibility.

The buried fact made the lesson measurable. Reading company documents was not an abstract virtue or a checklist feature. In this case, it separated a full-price €55,000 win from an automatic loss. For anyone accustomed to judging resilience by what happens when normal conditions fail, that is the result worth watching.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Apple releasing 20th anniversary iPhone, AirPods with cameras next year: report

Apple is expected to release a special 20th anniversary iPhone and new AirPods with cameras next year, according to reports from 9to5Mac. Details are still emerging.

Rust Glancer: Rust LSP Using 100X Less RAM

Rust Glancer, an experimental Rust language server, reduces RAM usage by 100 times compared to existing solutions, promising more efficient development tools.

Apple’s iPhone 18 Pro Features: Everything We Know So Far

A detailed overview of confirmed and rumored features of Apple’s upcoming iPhone 18 Pro, including design, camera, and performance updates.

Amazon Prime Day Isn’t the Only Major Event Happening This Week — Here Are 14 Can’t-Miss Sales

Several significant shopping events are happening this week, including 14 major sales across various retailers, alongside Amazon Prime Day. Here’s what you need to know.