firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

Good judgment begins with looking deeper

In natural wellness, the most visible ingredient is not always the one that determines the quality of a blend. Context matters: what else is present, how the elements interact and whether someone has examined the full picture before making a recommendation.

The same principle now applies to artificial intelligence in business. An AI agent may recognize a problem, communicate clearly and resist pressure, yet still fail because it did not investigate the company’s own records deeply enough. Firmulate’s live experiment turned that distinction into something measurable—and attached a €55,000 consequence to it.

The decisive information was not in the customer event placed directly before the models. It was buried two document references deep in the software company’s own files. The models that found it won the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost automatically.

B00P1RP5Z2

Amazon Product B00P1RP5Z2

As an affiliate, we earn on qualifying purchases.

A test of whether AI actually does its homework

Firmulate runs frontier AI models as complete companies through real crises, money mechanics and temptations. In the Crucible League experiment, each model managed the same small software company through its worst week. The customers, emergencies and opportunities to cut corners were identical. Every decision was versioned and auditable.

The company itself is synthetic but operationally demanding: it has 13 synthetic employees, burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is watchable at firmulate.com/live.

The striking result was not that some models understood the customer and others did not. All models spotted every crisis. All refused every manipulation attempt. Yet only two signed the €55,000 deal their analysis had earned. Firmulate summarized the gap neatly: “Same diagnosis, same pitch — no signature.”

That makes the buried document more than a scavenger-hunt detail. It reveals a practical property that buyers can evaluate: does an AI read the available files before answering? In this case, research depth directly changed the commercial outcome. A convincing response was insufficient without the evidence needed to support and close the deal.

The final standings

The final Crucible League table for July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. A do-nothing baseline scored 26 because partial progress still counted.

The complete results and plain-language findings are available on the Firmulate benchmarks page. One comparison requires care: Kimi K3 ran with the API default and without an effort parameter, while the other participants ran at xhigh.

The benchmark also treats trust as a hard boundary. A single breach caps the total because, in Firmulate’s words, “no amount of good work outweighs a breach of trust.” Every model cleared the experiment’s social-engineering challenges. Fake CEO messages escalated across three stages, and a reporter tried to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning directly: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as completion

Opus 4.8 illustrates why this type of evaluation matters. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, but it finished last. The close remained on the table, while its operating discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four other models, though less strongly.

This is an uncomfortable but useful finding. Detailed reasoning can look impressive while leaving the essential action unfinished. For companies evaluating agents, the meaningful questions extend beyond whether a model can identify a crisis or draft a polished answer. They include whether it follows references, respects boundaries, escalates correctly and completes the work its analysis recommends.

Firmulate also exposes 242 real, unedited management decisions through a “guess the model” quiz at firmulate.com/quiz.html. For enterprises wanting a more specific test, its pilot applies the same wargame to a read-only export of their own business. Nothing writes back to real systems; information is available at firmulate.com/pilot.html and contact@firmulate.com.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What buyers should measure before hiring an AI workforce

The lesson is not that every AI failure will be hidden two references deep. It is that seemingly small research habits can decide whether an agent produces useful work or merely plausible language.

For wellness businesses, that distinction may feel familiar. Responsible recommendations depend on examining the relevant context instead of reacting to the most visible detail. Business agents deserve the same scrutiny. Before trusting one with a customer relationship, support queue, forecast or commercial opportunity, test whether it reads what is available, preserves trust under pressure and finishes what it starts.

Firmulate’s €55,000 deal turned “does its homework” from a marketing promise into an observable business outcome. The models shared the diagnosis and the pitch. Only the agents that followed the evidence far enough secured the result.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


You May Also Like

Blending Essential Oils: Art and Science

Meta Description: Master the art and science of blending essential oils to create harmonious scents that delight your senses and elevate your aromatherapy journey.

Hands-On Healing: The Aromatherapy Massage Hack That’ll Wow Your Clients!

You’ll discover how to elevate your massage sessions with essential oils that not only soothe but also transform your clients’ experience.

Apply, Breathe, Relax: Aromatherapy Application Secrets the Pros Keep Quiet!

Just when you think you know everything about aromatherapy, discover the hidden secrets that can elevate your relaxation experience to new heights.