firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Crypto markets reward speed, but a fast answer is not the same as a sound decision. Before trusting an AI with a trading desk, treasury or customer operation, ask a tougher question: can it spot danger, resist pressure and follow through when the stakes collide? Firmulate’s live company experiment puts that question to work.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A rough week, the same rules

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The experiment tests management behavior rather than polished chat responses.

GPT-5.6-Sol finished first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The league’s trust rule is blunt: a single breach of trust caps the total; “no amount of good work outweighs a breach of trust.”

Recognizing a crisis is only half the job

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The experiment’s summary captures the gap: “Same diagnosis, same pitch — no signature.” In a business setting, sound analysis matters only if it carries through to a decision.

The winning detail was hidden in the company’s own files, two document references deep. It was not in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. That is a revealing kind of test for crypto firms, where the decisive information may sit in customer history, internal policies or operational records rather than the most visible signal.

There was also a direct test of authority and trust. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness does not guarantee execution

Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but it finished last. It left the deal unsigned and discipline slipped: it tried to write into a locked department instead of escalating. The same weakness appeared, in weaker form, in all four models.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The results are a snapshot of this experiment, not a universal verdict on every model or every business.

From watching to testing your own business

The live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, with a public cash countdown. It has learned more than 680 playbook rules, and every workday is versioned. Readers can watch the experiment at Firmulate; 242 real, unedited management decisions also power a “guess the model” quiz.

For an enterprise, the next step is to test scenarios against its own business. Firmulate’s pilot uses a read-only export to create a digital twin, then runs crisis scenarios and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. For a crypto business, that means exploring how an AI might handle a volatile week or a suspicious request without handing it the keys to live operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

The experiment shows why AI readiness is about more than spotting trouble: models also have to find the relevant evidence, respect boundaries and complete the decision. Enterprises can run the wargame against a read-only export of their own business. Explore the Firmulate pilot or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Death of the Identical Paragraph

A historic shift in news distribution as the traditional wire service model erodes due to AI-driven rewriting costs, raising questions about attribution and funding.

A Developer’s Guide To Selecting AI Models For Code Writing

This guide explains how developers can choose the right AI models for different coding tasks, improving efficiency and accuracy in software development.

Transform Your Content Workflow With 12 Top AI Tools In 2026

Explore 12 leading AI tools shaping content creation in 2026, streamlining research, drafting, publishing, and moderation for diverse channels.

Why Does Portland Have Such Long Summer Days? Applied Science Explains

Applied science clarifies why Portland has nearly 15 hours of daylight during summer solstice, revealing the astronomical factors behind extended daylight hours.