
Get audio and creator gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
When the tour is at risk, a polished demo won’t book the replacement venue
A creator business can turn on a single fragile chain: a distributor issue, a sponsor under pressure, a customer asking for a refund, or a message that looks like it came from the CEO. An AI assistant may sound confident about all of it. The harder question is whether it can make the right call when the stakes meet, and then follow through.
Firmulate puts that question into a live company experiment. The public brand runs AI models as companies facing real money mechanics, crises and temptations. Now it is inviting enterprises to take the idea from watching to testing their own business.
Same bad week, different models
In the final Crucible League, dated July 2026, each frontier model ran the same small software company through its worst week: same customers, same crises and same temptations. Decisions were versioned and auditable, so the contest was about what each model did, not how persuasive its chat sounded.
The results put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s integrity rule puts a sharp limit on a breach of trust: “no amount of good work outweighs a breach of trust.”
The models all spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding was succinct: “Same diagnosis, same pitch — no signature.” For a label, studio or creator platform considering AI agents, that gap matters. Recognizing the right move is different from carrying it through.
The detail hidden in the files
The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The result is a reminder that useful judgment can depend on finding evidence scattered across ordinary business records.
The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
A live company, not a staged chat
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. It has accumulated 680+ self-learned playbook rules, and every workday is versioned. Readers can watch the experiment at firmulate.com.
One participant shows why overall capability can hide operational weakness. Opus 4.8 was the most thorough, adding +80 learned rules and producing the deepest analyses, yet finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
There is a qualification to the ranking: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html.
From watching to a business pilot
The enterprise offer is to run the same kind of wargame against a read-only export of a company’s own business. The export can represent customers, pipeline and rules; scenarios can include churn, price increases, competitor attacks, PR crises and social-engineering pressure. The output is a board report with model rankings and weaknesses exposed in the company’s own playbooks.
That gives a music or creator business a way to examine how an AI workforce handles its own pressure points: for example, a disrupted release, a sponsor dispute or a rush to respond to a public complaint. Those are possible scenarios, not reported outcomes from the league. The stated boundary is clear: nothing writes back to real systems.

Test the judgment before handing over the keys
The league suggests that refusal and crisis recognition are only part of the job. Finding buried evidence, escalating when blocked and actually closing a deal also matter. A pilot lets an enterprise examine those behaviors against its own business context before deploying agents in workflows such as customer support, sales or operations.
To discuss a pilot using a read-only business export, visit Firmulate’s pilot page or email contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
