📊 Full opportunity report: The AI Competition Continues After The Demo – Here’s Why It Matters on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
The ongoing AI competition tests models in managing a live company under crisis conditions. Recent results show models excel at diagnosis but struggle with execution and trust, highlighting new evaluation needs.
Recent results from the Firmulate AI management experiment show that frontier models can identify crises and refuse manipulation but often fail to complete critical commercial tasks or maintain trust. This ongoing competition reveals important insights into how AI models perform in complex, real-world management scenarios, emphasizing the need for new evaluation standards.
In the latest July 2026 Crucible League, five AI models competed in managing a simulated small business facing crises, with gpt-5.6-sol leading at 95 points. The experiment measured not only diagnostic accuracy but also trustworthiness, decision-making, and execution, with a strict cap on trust breaches.
Despite all models correctly identifying crises and resisting manipulation attempts—such as fake CEO messages—they often failed to finalize deals or escalate issues properly. For example, only two models signed a €55,000 deal, even though their analysis identified the opportunity. The most thorough model, Opus 4.8, produced detailed analyses but failed to escalate or complete tasks effectively, finishing last overall.
The results underscore that high-quality responses alone do not guarantee successful management. The models’ ability to read organizational files, prioritize actions, and maintain trust under pressure remains inconsistent. Contextual factors, such as default API settings and effort parameters, also influenced outcomes, adding complexity to the evaluation.
The AI Competition Continues After the Demo — Here’s Why It Matters
The ongoing AI competition tests frontier models in managing a live company under crisis conditions. Recent results show models excel at diagnosis but struggle with execution and trust — revealing what traditional benchmarks never measure.
What the Firmulate Experiment Actually Measures
Crisis Identification
Models must read organizational files, spot the company’s worst-week crisis, and prioritize actions under pressure — beyond mere chat competence.
Trustworthiness Under Fire
With real money mechanics and versioned decision logs, models face manipulation attempts — including fake CEO messages — with a strict cap on trust breaches.
Decision Execution
Diagnosis alone doesn’t score. Models must finalize commercial deals, escalate critical issues, and complete management tasks end-to-end.
Diagnosis Is Solved. Execution Is Not.
The most thorough model in the league produced the most detailed analyses — yet finished last overall. It failed to escalate issues or complete tasks effectively, proving that high-quality responses alone do not guarantee successful management.
Where Models Excel — and Where They Stall
| Management Task | Model Performance | Why It Matters |
|---|---|---|
| Crisis diagnosis | ✓ All 5 models correct | Frontier models reliably read the situation — the easy part. |
| Resisting manipulation (fake CEO messages) | ✓ Refused across the board | Trust boundaries hold — even under simulated social engineering. |
| Closing the €55,000 deal | ✗ Only 2 of 5 signed | Analysis identified the opportunity; execution failed to follow through. |
| Escalating critical issues | ~ Inconsistent | Models often left problems unraised, a serious risk in real firms. |
| Reading organizational files accurately | ~ Inconsistent | Prioritization depends on context the models sometimes miss. |
Trust and Execution Are Separate Challenges
“The key insight is that high-quality answers are not enough; effective management requires trust, prioritization, and the ability to complete the job under pressure.”
— Thorsten Meyer, Lead Researcher“Our models refused manipulation attempts but still failed to escalate critical issues, showing that trust and execution are separate challenges.”
— A Participating AI DeveloperFrom Diagnosis to Deployment Readiness
🔍 Diagnose
Models read files and identify the crisis in the simulated company’s worst week.
🛡️ Hold Trust
Resist manipulation attempts and stay within the trust-breach cap.
⚡ Execute
Finalize deals, escalate issues, and complete commercial tasks.
📊 Score & Deploy
Ongoing benchmarks assess readiness for real business management roles.
What This Means for AI in Business
What does the Firmulate experiment measure?
The experiment assesses AI models’ ability to diagnose crises, prioritize actions, maintain trust, and complete management tasks in a simulated business environment.
Why are trust and execution important in AI management?
Trust ensures that models do not bypass ethical or security boundaries, while effective execution determines whether they can actually accomplish business objectives under real-world conditions.
What are the main weaknesses revealed by the latest results?
Models often fail to escalate issues properly, complete commercial deals, or read organizational files accurately — despite identifying crises and resisting manipulation.
Will these benchmarks influence how AI is deployed in companies?
Yes. Organizations will likely prioritize models that demonstrate strong management skills — trustworthiness and task completion — before integrating AI into critical workflows.
Future evaluations will incorporate longer-term management tasks, broader organizational contexts, and real-time decision consequences — possibly including live business deployments under controlled conditions. Whether results generalize beyond the simulation, and how API settings and effort parameters shape outcomes, remain open questions.
Implications for AI Management Evaluation
This ongoing competition highlights that AI models’ management skills—including trust, decision execution, and prioritization—are crucial metrics beyond mere response quality. As AI tools become integrated into business operations, understanding their real-world management capabilities is vital for safe and effective deployment. The results suggest that future benchmarks must assess not just what models say but how they manage consequences, escalate issues, and uphold trust over time.
As an affiliate, we earn on qualifying purchases.
Background of the Firmulate Experiment
Launched as a live test of frontier AI models’ management skills, the Firmulate experiment simulates a small company’s worst week, with real money mechanics and versioned decision logs. The goal is to evaluate models’ ability to diagnose crises, prioritize actions, and maintain organizational trust, offering a more realistic assessment than traditional benchmarks like coding or chat competitions.
Previous results showed models could identify crises but struggled with execution and trust, prompting ongoing refinement of evaluation criteria. The July 2026 league introduced stricter standards, including trust breaches and real-world decision-making challenges, to deepen insights into AI readiness for business management roles.
“The key insight is that high-quality answers are not enough; effective management requires trust, prioritization, and the ability to complete the job under pressure.”
— Thorsten Meyer, Lead Researcher
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Uncertainties in Model Performance
It is still unclear how much the models’ performance varies across different business scenarios or industries. The impact of API settings, effort parameters, and organizational complexity on outcomes also requires further investigation. Additionally, whether these results generalize beyond the simulated environment remains to be seen, as real companies may present even greater unpredictability.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Benchmarks
Future evaluations will likely incorporate longer-term management tasks, broader organizational contexts, and real-time decision consequences. Developers and organizations will need to refine their models to improve trust, escalation, and task completion, with ongoing benchmarks serving as critical tools for assessing progress. The next phase may also involve deploying models in live business environments under controlled conditions to validate their capabilities further.
AI trustworthiness evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the Firmulate experiment measure?
The experiment assesses AI models’ ability to diagnose crises, prioritize actions, maintain trust, and complete management tasks in a simulated business environment.
Why are trust and execution important in AI management?
Trust ensures that AI models do not bypass ethical or security boundaries, while effective execution determines whether the models can actually accomplish business objectives under real-world conditions.
What are the main weaknesses revealed by the latest results?
Models often fail to escalate issues properly, complete commercial deals, or read organizational files accurately, despite identifying crises and resisting manipulation.
Will these benchmarks influence how AI is deployed in companies?
Yes, organizations will likely prioritize models that demonstrate strong management skills, including trustworthiness and task completion, before integrating AI into critical workflows.
What is the significance of the different scores among models?
The scores reflect not only diagnostic accuracy but also the models’ ability to manage trust, escalate issues, and complete tasks, revealing gaps in current AI capabilities for management roles.
Source: ThorstenMeyerAI.com