🔍 Read the full analysis: What To Test Before AI Agents Start Handling Business Tasks on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Firmulate says its final Crucible League ran frontier models through a simulated software company’s difficult week, scoring crisis response, judgment and execution. The results showed a gap between recognizing a crisis and closing a justified deal; the company’s enterprise pilot proposes testing scenarios against read-only business data.
Firmulate says its final Crucible League, completed in July 2026, put frontier AI models through a simulated software company’s hardest week, testing whether they could identify crises, protect trust and act on available evidence. The company’s next step is an enterprise pilot using a read-only export of a client’s business data, intended to assess agent behavior without writing to live systems.
The league ranked gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73; a do-nothing baseline scored 26. Firmulate says the scoring allowed partial progress but capped a participant’s total after a breach of trust. It described that rule as: “no amount of good work outweighs a breach of trust.” These are results from Firmulate’s experiment, not an independent evaluation of general model performance.
According to Firmulate, all five models spotted every crisis and refused every manipulation attempt. The gap appeared in follow-through: only two signed a €55,000 deal that their own analysis supported. The competitor weakness needed to make the case was buried two document references deep in the company files. Models that found it closed the deal at full price, which Firmulate valued at €4,583 in monthly recurring revenue.
The trust tests included staged fake CEO messages and a reporter asking for a yes-or-no answer “on background.” Firmulate reports that all five models refused. It also says Opus 4.8 generated the most learned rules, adding 80, and produced the deepest analyses, but still finished last. Among its errors was an attempt to write into a locked department rather than escalate. Firmulate says a weaker version of that boundary problem appeared in all four other models.
What To Test Before AI Agents Start Handling Business Tasks
Firmulate’s final Crucible League ran five frontier AI models through a simulated software company’s hardest week — testing crisis recognition, trust under manipulation, and the ability to follow through on justified decisions. The enterprise pilot proposes the same scenarios against read-only exports of real business data.
“No amount of good work outweighs a breach of trust.”
Firmulate — Scoring RuleThe Standings: Recognition vs. Execution
Firmulate’s scoring allowed partial progress but capped a participant’s total after a breach of trust. All five models passed the crisis and manipulation tests — the spread emerged in evidence-finding and deal-closing.
The Capability Gaps That Matter
An agent can identify an emergency and resist impersonation while still missing buried evidence, failing to close a justified deal, or mishandling a blocked permission. These gaps matter in customer support, sales and internal operations.
Spotting the Emergency
All five models identified every crisis in the simulated week. Recognition is a solved problem in this test — but it is only the entry barrier, not the differentiator.
Refusing the Bypass
Staged fake CEO messages and a reporter seeking a yes-or-no answer “on background” were refused by all five models — including Opus 4.8, which still finished last overall.
The Follow-Through Gap
The competitor weakness justifying the deal was buried two document references deep. Models that found it closed at full price; only two of five did.
Testing Against Real Business Data — Read-Only
Firmulate’s next step shifts the exercise from a synthetic company to a client’s own information: crisis scenarios built on customers, pipeline, rules and pressure points, delivered as a board report.
The synthetic company runs on just €2,300 in monthly recurring revenue, with a public cash countdown and versioned workdays — a deliberately fragile testbed.
Locked Departments & Escalation
Opus 4.8 tried to write into a locked department rather than escalate — generating 80 new learned rules and the deepest analyses, yet still finishing last. A weaker form of the same boundary problem appeared in all four other models.
How a Crucible-Style Evaluation Runs
The chain from synthetic simulation to board-level insight, as Firmulate describes its methodology and proposed pilot.
Simulation Week
Frontier models run a synthetic software company through its hardest week — crises, pressure and staged manipulations.
Trust-Capped Scoring
Partial progress counts, but any breach of trust caps the total — regardless of how strong the rest of the work is.
Read-Only Client Export
The enterprise pilot replays scenarios against a client’s own customers, pipeline and rules — with nothing written back.
Board Report
Model rankings plus weak points in the company’s playbooks, highlighting where agents fail under real conditions.
What the Agents Actually Said
“Same diagnosis, same pitch — no signature.”
Firmulate“Treat the request as a suspected approval-bypass / possible impersonation.”
Kimi K3, as quoted by Firmulate“No amount of good work outweighs a breach of trust.”
Firmulate — Scoring RuleWhat the Results Do — and Don’t — Tell You
The published account leaves open questions a company evaluating such a test should ask before drawing conclusions for its own operations.
| Question | Published? | What We Know |
|---|---|---|
| Scoring methodology detail | ~ Partial | Trust-cap rule described; full scoring method not specified enough to judge predictive value. |
| All five models detected crises & refused manipulation | ✓ Yes | Every model passed every crisis and manipulation test in the simulated week. |
| Identity of the two deal-closing models | ✗ No | The published account does not identify which models signed the €55,000 deal. |
| Equal effort settings across models | ✗ No | Kimi K3 ran at API default; the others ran at xhigh — direct comparisons are complicated. |
| Pilot writes back to live systems | ✓ No Write-Back | Enterprise pilot uses a read-only export; nothing writes to real systems. |
| Data handling, retention & access controls for pilot | ✗ No | Not detailed — nor pilot pricing, timetables, or independent validation process. |
| Predictive of performance in other companies | ✗ No | Tests one designed simulated week; no evidence of transfer to unscripted real-world work. |
From Crisis Recognition to Follow-Through
The results highlight distinct capabilities that businesses may need to evaluate before giving agents operational tasks. An agent can identify an emergency and resist an apparent impersonation attempt while still missing evidence in company records, failing to complete a justified commercial action or mishandling a blocked permission. Those gaps can matter in customer support, sales and internal operations, where the quality of the decision depends on both context and execution.
Firmulate’s proposed pilot shifts the exercise from a synthetic company to a client’s own information. Its stated aim is to test scenarios involving a company’s customers, pipeline, rules and pressure points, then provide a board report with model rankings and weaknesses in its playbooks. A read-only setup limits the test’s direct effect on business systems, though it does not by itself establish how representative the scenarios or scores are of live work.
A Simulated Company Under Pressure
Firmulate’s live experiment follows a synthetic company with 13 employees, a stated monthly burn of €105,000 and €2,300 in monthly recurring revenue. The company also reports a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Visitors can follow the simulation and take a quiz based on 242 real, unedited management decisions, guessing which model made each choice.
The league’s ranking needs one qualification: Firmulate says Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference is part of the experiment’s conditions, and the published scores should be read in that context. The experiment tests a defined simulated week; it does not establish how the same models would perform across companies or unscripted real-world situations.
“Same diagnosis, same pitch — no signature.”
— Firmulate
Limits of the Published Results
The published account does not specify enough about the scoring method, scenario design or independent review to determine how predictive the league is of performance in a real company. It also does not identify the two models that signed the deal in the account of the result. The difference in Kimi K3’s effort setting further complicates direct comparisons across the standings.
For the enterprise pilot, Firmulate says the export is read-only and that nothing writes back to real systems. The published description does not detail data handling, retention, access controls, pilot pricing or how scenarios and rankings would be tailored for each client. Those details would be relevant to a company evaluating a test with its own business records.
Company-Specific Pilots Ahead
Firmulate invites companies to discuss a pilot based on a read-only export. It says the exercise would run crisis scenarios and produce a board report covering model rankings and weak points in a company’s playbooks. No timetable, participant list or independent validation process for these pilots is specified in the published account.
Readers can follow the synthetic company at firmulate.com/live and review the league results at firmulate.com/benchmarks.html. The broader business question remains whether a model that performs well in a designed wargame can reliably find relevant evidence, complete authorized work and respect boundaries under the conditions of a particular company.
Source: ThorstenMeyerAI.com
Key Questions
What did Firmulate’s Crucible League test?
It put five frontier models through a simulated software company’s difficult week, testing crisis recognition, trust and follow-through on business decisions.
Which model ranked first?
Firmulate’s standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93. The results come from this experiment, and Kimi K3 used a different effort setting from the other models.
What does the enterprise pilot involve?
Firmulate says it uses a read-only export of a company’s data to run crisis scenarios and prepare a board report on model rankings and playbook weaknesses. It says the pilot does not write back to real systems.
Do the results show how agents will perform in a live business?
No. They describe performance in Firmulate’s specific simulation. The published account does not establish how predictive those scores are of performance in other companies or live operations.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
