
Good judgment begins with looking deeper
In natural wellness, the most visible ingredient is not always the one that determines the quality of a blend. Context matters: what else is present, how the elements interact and whether someone has examined the full picture before making a recommendation.
The same principle now applies to artificial intelligence in business. An AI agent may recognize a problem, communicate clearly and resist pressure, yet still fail because it did not investigate the company’s own records deeply enough. Firmulate’s live experiment turned that distinction into something measurable—and attached a €55,000 consequence to it.
The decisive information was not in the customer event placed directly before the models. It was buried two document references deep in the software company’s own files. The models that found it won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost automatically.
A test of whether AI actually does its homework
Firmulate runs frontier AI models as complete companies through real crises, money mechanics and temptations. In the Crucible League experiment, each model managed the same small software company through its worst week. The customers, emergencies and opportunities to cut corners were identical. Every decision was versioned and auditable.
The company itself is synthetic but operationally demanding: it has 13 synthetic employees, burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable at firmulate.com/live.
The striking result was not that some models understood the customer and others did not. All models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. Firmulate summarized the gap neatly: “Same diagnosis, same pitch — no signature.”
That makes the buried document more than a scavenger-hunt detail. It reveals a practical property that buyers can evaluate: does an AI read the available files before answering? In this case, research depth directly changed the commercial outcome. A convincing response was insufficient without the evidence needed to support and close the deal.
The final standings
The final Crucible League table for July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress still counted.
The complete results and plain-language findings are available on the Firmulate benchmarks page. One comparison requires care: Kimi K3 ran with the API default and without an effort parameter, while the other participants ran at xhigh.
The benchmark also treats trust as a hard boundary. A single breach caps the total because, in Firmulate’s words, “no amount of good work outweighs a breach of trust.” Every model cleared the experiment’s social-engineering challenges. Fake CEO messages escalated across three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as completion
Opus 4.8 illustrates why this type of evaluation matters. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, but it finished last. The close remained on the table, while its operating discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.
This is an uncomfortable but useful finding. Detailed reasoning can look impressive while leaving the essential action unfinished. For companies evaluating agents, the meaningful questions extend beyond whether a model can identify a crisis or draft a polished answer. They include whether it follows references, respects boundaries, escalates correctly and completes the work its analysis recommends.
Firmulate also exposes 242 real, unedited management decisions through a “guess the model” quiz at firmulate.com/quiz.html. For enterprises wanting a more specific test, its pilot applies the same wargame to a read-only export of their own business. Nothing writes back to real systems; information is available at firmulate.com/pilot.html and contact@firmulate.com.

What buyers should measure before hiring an AI workforce
The lesson is not that every AI failure will be hidden two references deep. It is that seemingly small research habits can decide whether an agent produces useful work or merely plausible language.
For wellness businesses, that distinction may feel familiar. Responsible recommendations depend on examining the relevant context instead of reacting to the most visible detail. Business agents deserve the same scrutiny. Before trusting one with a customer relationship, support queue, forecast or commercial opportunity, test whether it reads what is available, preserves trust under pressure and finishes what it starts.
Firmulate’s €55,000 deal turned “does its homework” from a marketing promise into an observable business outcome. The models shared the diagnosis and the pitch. Only the agents that followed the evidence far enough secured the result.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html