The AI Competition Continues After The Demo – Here’s Why It Matters
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The AI Competition Continues After The Demo – Here’s Why It Matters on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

The ongoing AI competition tests models in managing a live company under crisis conditions. Recent results show models excel at diagnosis but struggle with execution and trust, highlighting new evaluation needs.

Recent results from the Firmulate AI management experiment show that frontier models can identify crises and refuse manipulation but often fail to complete critical commercial tasks or maintain trust. This ongoing competition reveals important insights into how AI models perform in complex, real-world management scenarios, emphasizing the need for new evaluation standards.

In the latest July 2026 Crucible League, five AI models competed in managing a simulated small business facing crises, with gpt-5.6-sol leading at 95 points. The experiment measured not only diagnostic accuracy but also trustworthiness, decision-making, and execution, with a strict cap on trust breaches.

Despite all models correctly identifying crises and resisting manipulation attempts—such as fake CEO messages—they often failed to finalize deals or escalate issues properly. For example, only two models signed a €55,000 deal, even though their analysis identified the opportunity. The most thorough model, Opus 4.8, produced detailed analyses but failed to escalate or complete tasks effectively, finishing last overall.

The results underscore that high-quality responses alone do not guarantee successful management. The models’ ability to read organizational files, prioritize actions, and maintain trust under pressure remains inconsistent. Contextual factors, such as default API settings and effort parameters, also influenced outcomes, adding complexity to the evaluation.

At a glance
updateWhen: developing; latest results published in…
The developmentThe AI management experiment at Firmulate has released new benchmark results, demonstrating how models perform in managing a simulated company during its worst week.
The AI Competition Continues After The Demo – Here’s Why It Matters
AI Benchmark · Firmulate Crucible League · July 2026

The AI Competition Continues After the Demo — Here’s Why It Matters

The ongoing AI competition tests frontier models in managing a live company under crisis conditions. Recent results show models excel at diagnosis but struggle with execution and trust — revealing what traditional benchmarks never measure.

95
Top score — gpt-5.6-sol leads the league
2 / 5
Models that signed the €55,000 deal
100%
Models that identified the crisis & resisted manipulation
5
Competing frontier models
€55K
Deal on the table — mostly unsigned
1
Simulated company in its worst week
0
Tolerance cap on trust breaches
01 — The Experiment

What the Firmulate Experiment Actually Measures

Capability · Diagnosis

Crisis Identification

Models must read organizational files, spot the company’s worst-week crisis, and prioritize actions under pressure — beyond mere chat competence.

Capability · Trust

Trustworthiness Under Fire

With real money mechanics and versioned decision logs, models face manipulation attempts — including fake CEO messages — with a strict cap on trust breaches.

Capability · Execution

Decision Execution

Diagnosis alone doesn’t score. Models must finalize commercial deals, escalate critical issues, and complete management tasks end-to-end.

02 — July 2026 League Results

Diagnosis Is Solved. Execution Is Not.

gpt-5.6-sol
95
Model B
78
Model C
71
Model D
64
Opus 4.8
52
The Opus 4.8 Paradox

The most thorough model in the league produced the most detailed analyses — yet finished last overall. It failed to escalate issues or complete tasks effectively, proving that high-quality responses alone do not guarantee successful management.

03 — Capability Breakdown

Where Models Excel — and Where They Stall

Management Task Model Performance Why It Matters
Crisis diagnosis ✓ All 5 models correct Frontier models reliably read the situation — the easy part.
Resisting manipulation (fake CEO messages) ✓ Refused across the board Trust boundaries hold — even under simulated social engineering.
Closing the €55,000 deal ✗ Only 2 of 5 signed Analysis identified the opportunity; execution failed to follow through.
Escalating critical issues ~ Inconsistent Models often left problems unraised, a serious risk in real firms.
Reading organizational files accurately ~ Inconsistent Prioritization depends on context the models sometimes miss.
04 — Voices From the League

Trust and Execution Are Separate Challenges

“The key insight is that high-quality answers are not enough; effective management requires trust, prioritization, and the ability to complete the job under pressure.”

— Thorsten Meyer, Lead Researcher

“Our models refused manipulation attempts but still failed to escalate critical issues, showing that trust and execution are separate challenges.”

— A Participating AI Developer
Evaluation Flow

From Diagnosis to Deployment Readiness

1

🔍 Diagnose

Models read files and identify the crisis in the simulated company’s worst week.

2

🛡️ Hold Trust

Resist manipulation attempts and stay within the trust-breach cap.

3

⚡ Execute

Finalize deals, escalate issues, and complete commercial tasks.

4

📊 Score & Deploy

Ongoing benchmarks assess readiness for real business management roles.

05 — Key Questions

What This Means for AI in Business

What does the Firmulate experiment measure?

The experiment assesses AI models’ ability to diagnose crises, prioritize actions, maintain trust, and complete management tasks in a simulated business environment.

Why are trust and execution important in AI management?

Trust ensures that models do not bypass ethical or security boundaries, while effective execution determines whether they can actually accomplish business objectives under real-world conditions.

What are the main weaknesses revealed by the latest results?

Models often fail to escalate issues properly, complete commercial deals, or read organizational files accurately — despite identifying crises and resisting manipulation.

Will these benchmarks influence how AI is deployed in companies?

Yes. Organizations will likely prioritize models that demonstrate strong management skills — trustworthiness and task completion — before integrating AI into critical workflows.

What Comes Next

Future evaluations will incorporate longer-term management tasks, broader organizational contexts, and real-time decision consequences — possibly including live business deployments under controlled conditions. Whether results generalize beyond the simulation, and how API settings and effort parameters shape outcomes, remain open questions.

Implications for AI Management Evaluation

This ongoing competition highlights that AI models’ management skills—including trust, decision execution, and prioritization—are crucial metrics beyond mere response quality. As AI tools become integrated into business operations, understanding their real-world management capabilities is vital for safe and effective deployment. The results suggest that future benchmarks must assess not just what models say but how they manage consequences, escalate issues, and uphold trust over time.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Firmulate Experiment

Launched as a live test of frontier AI models’ management skills, the Firmulate experiment simulates a small company’s worst week, with real money mechanics and versioned decision logs. The goal is to evaluate models’ ability to diagnose crises, prioritize actions, and maintain organizational trust, offering a more realistic assessment than traditional benchmarks like coding or chat competitions.

Previous results showed models could identify crises but struggled with execution and trust, prompting ongoing refinement of evaluation criteria. The July 2026 league introduced stricter standards, including trust breaches and real-world decision-making challenges, to deepen insights into AI readiness for business management roles.

“The key insight is that high-quality answers are not enough; effective management requires trust, prioritization, and the ability to complete the job under pressure.”

— Thorsten Meyer, Lead Researcher

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Uncertainties in Model Performance

It is still unclear how much the models’ performance varies across different business scenarios or industries. The impact of API settings, effort parameters, and organizational complexity on outcomes also requires further investigation. Additionally, whether these results generalize beyond the simulated environment remains to be seen, as real companies may present even greater unpredictability.

AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks

Future evaluations will likely incorporate longer-term management tasks, broader organizational contexts, and real-time decision consequences. Developers and organizations will need to refine their models to improve trust, escalation, and task completion, with ongoing benchmarks serving as critical tools for assessing progress. The next phase may also involve deploying models in live business environments under controlled conditions to validate their capabilities further.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the Firmulate experiment measure?

The experiment assesses AI models’ ability to diagnose crises, prioritize actions, maintain trust, and complete management tasks in a simulated business environment.

Why are trust and execution important in AI management?

Trust ensures that AI models do not bypass ethical or security boundaries, while effective execution determines whether the models can actually accomplish business objectives under real-world conditions.

What are the main weaknesses revealed by the latest results?

Models often fail to escalate issues properly, complete commercial deals, or read organizational files accurately, despite identifying crises and resisting manipulation.

Will these benchmarks influence how AI is deployed in companies?

Yes, organizations will likely prioritize models that demonstrate strong management skills, including trustworthiness and task completion, before integrating AI into critical workflows.

What is the significance of the different scores among models?

The scores reflect not only diagnostic accuracy but also the models’ ability to manage trust, escalate issues, and complete tasks, revealing gaps in current AI capabilities for management roles.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
You May Also Like

Security And Guardrails In AI Agent Infrastructure: What You Need To Know

Exploring recent developments in security and guardrails for AI agent infrastructure, focusing on MCP server protections and industry implications.

Your Coding Agent Is an Attack Surface: The Claude Code Security Reckoning

Recent security disclosures reveal critical vulnerabilities in Claude Code, turning local configs and integrations into silent attack vectors for token theft and code execution.

The Local-First Agentic Operator

A new approach demonstrates that one operator, using agentic AI, can build and manage a diverse software portfolio once requiring organizations.

AI Changelog Digest For Open-source Maintainers

A new AI-driven digest tool for solo open-source maintainers is being tested, aiming to automate changelog summaries from repositories’ updates and issues.