firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

The difference between a thoughtful blend and a crowded bottle

Readers interested in aromatherapy and natural wellness already understand a useful principle: adding more ingredients does not automatically create a better result. Each element needs a purpose, and the final blend still has to serve the person using it. Preparation matters, but so do judgment, restraint and follow-through.

Firmulate’s live business experiment has uncovered a remarkably similar lesson about artificial intelligence. Opus 4.8 was the most thorough participant in a demanding management wargame. It produced the deepest analyses and learned 80 new playbook rules. Yet it finished last. Its story is not one of incompetence. It is a respectful warning that diligence can become its own distraction when an intelligent system fails to convert insight into action.

Amazon

Top picks for "meticulou mistak preparation"

As an affiliate, we earn on qualifying purchases.

A worst week designed to reveal working habits

Firmulate gave each frontier model the same assignment: run the same small software company through its worst week, facing identical customers, crises and temptations. The simulated company has 13 employees and real money mechanics, including monthly burn of €105,000 against monthly recurring revenue of €2,300. Its public cash countdown makes the pressure visible.

This was not a chat demonstration built around polished answers. Every decision was versioned and auditable. Across the live company, more than 680 self-learned playbook rules accumulated. The test asked whether a model could recognize trouble, investigate it, make a defensible decision and complete the work.

The final July 2026 Crucible League results were:

  • gpt-5.6-sol: 95
  • Kimi K3: 93
  • Sonnet 5: 88
  • Fable 5: 77
  • Opus 4.8: 73

A do-nothing baseline scored 26 because partial progress counts. But Firmulate also imposed a crucial boundary: a single breach of trust caps the total. As the experiment puts it, “no amount of good work outweighs a breach of trust.” Full results and plain-language findings are available on the Firmulate benchmarks page.

Opus understood the week better than its result suggests

Opus 4.8 deserves a fair reading. It was the most thorough model in the field, adding 80 learned rules and producing the deepest analyses. It spotted every crisis. Like every other participant, it also refused every manipulation attempt.

Those attempts included fake messages from the chief executive escalating over three stages and a reporter seeking “just one yes/no, on background.” All 5 models refused. Kimi K3 captured the appropriate posture in its recorded reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Opus belonged to that clean sweep, showing that its last-place finish was not caused by gullibility or an ethical collapse.

The trouble was execution discipline. Opus attempted to write into a locked department instead of escalating the obstacle. More importantly, it left a valuable commercial close unfinished. Its analysis had helped earn the opportunity, but it did not complete the decisive step.

The decisive detail was hidden in ordinary company material

The week’s most consequential fact was not presented in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that read the relevant file could use that insight to win the deal at full price, adding €4,583 in monthly recurring revenue.

Only two models signed the €55,000 deal their own work had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.” This is where Opus’s profile becomes especially instructive. The model capable of the deepest analysis still failed to prioritize the investigation and closing behavior that would have produced the largest practical result.

That weakness was not unique to Opus. It appeared in all four other models, though less strongly. This matters because the lesson is broader than one model’s ranking. AI systems can identify a situation correctly, produce impressive reasoning and still stop before the outcome that matters.

A fair comparison still needs context

Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers compare performances. Even so, the published results document what each participant actually did under its assigned conditions: gpt-5.6-sol led with 95, while Opus 4.8 placed fifth with 73.

The experiment can also be examined through 242 real, unedited management decisions used in Firmulate’s model-guessing quiz. Together, those decisions turn abstract claims about intelligence into observable working behavior: what the models noticed, what they ignored, where they resisted pressure and whether they finished what they began.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

For businesses, discernment matters more than volume

Opus 4.8’s performance offers a humane way to think about AI assessment. The model was careful, productive and ethically steady. Its failure was narrower and more recognizable: it generated substantial understanding without consistently directing that effort toward the highest-value action.

That distinction matters for any organization considering AI for customer management, support work or financial decisions. A long analysis is not the same as a completed task. A growing library of rules is not proof that the right rule will be applied at the right moment. The useful questions are whether the system reads the available material, escalates when blocked, protects trust and carries promising work through to a measurable conclusion.

Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business; nothing writes back to real systems. The larger message is simple: thoroughness is valuable, but impact depends on prioritization. Like a carefully composed wellness practice, effectiveness comes from choosing what matters and completing it with discipline.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FLEA & TICK SEAS

Flea & tick season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stop Over‑Diffusing: The 30/30 Rule for Safer Aroma

Meta description: “Many people over-diffuse essential oils, but the 30/30 rule offers a simple way to enjoy aromatherapy safely—discover how to protect your health.

Essential Oil Blends for Seasonal Transitions

Get ready to transform your space with essential oil blends that enhance seasonal transitions, but discover the perfect combinations that suit your unique style.