firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get oils, diffusers and self-care delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Before an AI handles the next customer crisis, give it a rehearsal

A natural wellness business may depend on customer trust as much as a carefully chosen oil or a reliable supplier. Imagine an AI assistant handling a sudden wave of cancellations, a competitor’s offer and a message pretending to come from the CEO. A polished answer in a chat window cannot show whether it will make the right decisions when those pressures arrive together.

That is the question behind Firmulate: can an AI workforce manage a company through a bad week, and where does its judgment falter? Its live experiment is watchable at firmulate.com. The next step is to rehearse those decisions against your own business.

The same company, the same hard week

In the final Crucible League, published in July 2026, frontier models faced the same small software company, customers, crises and temptations. The experiment tracked decisions so they could be reviewed. The final standings were gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

On one important measure, the models performed consistently: every model spotted every crisis and refused every manipulation attempt. But recognition was not the same as execution. Only two signed a €55,000 deal their own analysis had earned. The experiment’s concise description was: “Same diagnosis, same pitch — no signature.”

The answer was buried in the company’s own files

The decisive competitor weakness was not in a customer event. It sat two document references deep in the company’s files. Models that read the file won the deal at full price, worth +€4,583 MRR. The result points to a practical challenge for businesses: useful evidence may already exist in internal documents, but an AI has to find it and carry its implications through to a decision.

The pressure tests also included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

More analysis did not guarantee a better finish

Opus 4.8 was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That context matters when readers interpret a ranking. Firmulate also offers a quiz built from 242 real, unedited management decisions: visitors can guess which model made each decision at firmulate.com.

A rehearsal with real business stakes, kept out of live systems

The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. Its playbook has grown to 680+ self-learned rules, and every workday is versioned. These details make the experiment watchable as an ongoing company story, while keeping its employees synthetic.

For a business considering AI agents in customer support, sales or forecasting, the pilot moves from observation to a test of its own decisions. Firmulate says an enterprise can supply a read-only export of its business, then run crisis scenarios against that company data and receive a board report with a model ranking and weak points in its playbooks. Nothing writes back to real systems.

Test the judgment before handing over the work

The league suggests that spotting danger and refusing manipulation are only part of the job. Models also need to find relevant evidence, follow through on sound analysis and respect boundaries when they cannot act. A pilot can help a leadership team see how those choices play out against its own company’s information and scenarios.

To discuss a Firmulate enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Wellness content on this site is informational and not a substitute for professional medical guidance.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Natural First Aid: Essential Oil Emergency Kit Guide

Prepare your home with a natural first aid kit using essential oils for everyday ailments; discover the powerful benefits waiting for you inside!

Massage Mastery: Make Aromatherapy Oils for Massage That Clients Crave!

Master the art of crafting irresistible aromatherapy oils that enhance your massage practice and leave clients yearning for more. Discover the secrets within!

Essential Oils, Honest Scores, and the AI That Almost Ran a Company: Why a Do-Nothing Manager Gets 26 Points, Not 0

An AI benchmark where doing nothing scores 26, one breach of trust caps your grade, and only two of five models closed a deal they’d already earned.