The AI Competition Continues After The Demo – Here’s Why It Matters
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Before you orderOffer from Amazon

Get hardware and tech essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The ongoing AI competition tests models in managing a live company under crisis conditions. Recent results show models excel at diagnosis but struggle with execution and trust, highlighting new evaluation needs.

Recent results from the Firmulate AI management experiment show that frontier models can identify crises and refuse manipulation but often fail to complete critical commercial tasks or maintain trust. This ongoing competition reveals important insights into how AI models perform in complex, real-world management scenarios, emphasizing the need for new evaluation standards.

In the latest July 2026 Crucible League, five AI models competed in managing a simulated small business facing crises, with gpt-5.6-sol leading at 95 points. The experiment measured not only diagnostic accuracy but also trustworthiness, decision-making, and execution, with a strict cap on trust breaches.

Despite all models correctly identifying crises and resisting manipulation attempts—such as fake CEO messages—they often failed to finalize deals or escalate issues properly. For example, only two models signed a €55,000 deal, even though their analysis identified the opportunity. The most thorough model, Opus 4.8, produced detailed analyses but failed to escalate or complete tasks effectively, finishing last overall.

The results underscore that high-quality responses alone do not guarantee successful management. The models’ ability to read organizational files, prioritize actions, and maintain trust under pressure remains inconsistent. Contextual factors, such as default API settings and effort parameters, also influenced outcomes, adding complexity to the evaluation.

At a glance
updateWhen: developing; latest results published in…
The developmentThe AI management experiment at Firmulate has released new benchmark results, demonstrating how models perform in managing a simulated company during its worst week.

Implications for AI Management Evaluation

This ongoing competition highlights that AI models’ management skills—including trust, decision execution, and prioritization—are crucial metrics beyond mere response quality. As AI tools become integrated into business operations, understanding their real-world management capabilities is vital for safe and effective deployment. The results suggest that future benchmarks must assess not just what models say but how they manage consequences, escalate issues, and uphold trust over time.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of the Firmulate Experiment

Launched as a live test of frontier AI models’ management skills, the Firmulate experiment simulates a small company’s worst week, with real money mechanics and versioned decision logs. The goal is to evaluate models’ ability to diagnose crises, prioritize actions, and maintain organizational trust, offering a more realistic assessment than traditional benchmarks like coding or chat competitions.

Previous results showed models could identify crises but struggled with execution and trust, prompting ongoing refinement of evaluation criteria. The July 2026 league introduced stricter standards, including trust breaches and real-world decision-making challenges, to deepen insights into AI readiness for business management roles.

“The key insight is that high-quality answers are not enough; effective management requires trust, prioritization, and the ability to complete the job under pressure.”

— Thorsten Meyer, Lead Researcher

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Uncertainties in Model Performance

It is still unclear how much the models’ performance varies across different business scenarios or industries. The impact of API settings, effort parameters, and organizational complexity on outcomes also requires further investigation. Additionally, whether these results generalize beyond the simulated environment remains to be seen, as real companies may present even greater unpredictability.

Amazon

AI decision-making software for small business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks

Future evaluations will likely incorporate longer-term management tasks, broader organizational contexts, and real-time decision consequences. Developers and organizations will need to refine their models to improve trust, escalation, and task completion, with ongoing benchmarks serving as critical tools for assessing progress. The next phase may also involve deploying models in live business environments under controlled conditions to validate their capabilities further.

Amazon

trustworthy AI management platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the Firmulate experiment measure?

The experiment assesses AI models’ ability to diagnose crises, prioritize actions, maintain trust, and complete management tasks in a simulated business environment.

Why are trust and execution important in AI management?

Trust ensures that AI models do not bypass ethical or security boundaries, while effective execution determines whether the models can actually accomplish business objectives under real-world conditions.

What are the main weaknesses revealed by the latest results?

Models often fail to escalate issues properly, complete commercial deals, or read organizational files accurately, despite identifying crises and resisting manipulation.

Will these benchmarks influence how AI is deployed in companies?

Yes, organizations will likely prioritize models that demonstrate strong management skills, including trustworthiness and task completion, before integrating AI into critical workflows.

What is the significance of the different scores among models?

The scores reflect not only diagnostic accuracy but also the models’ ability to manage trust, escalate issues, and complete tasks, revealing gaps in current AI capabilities for management roles.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When a Content Network Starts Publishing to Itself

A large automated content network is publishing heavily to a few sites while neglecting others, revealing systemic issues in distribution logic.

What’s Behind The Buzz? Taco Bell’s Ice Cream Taco And Food Trends

Taco Bell’s new ice cream taco is gaining attention, highlighting emerging food trends and consumer curiosity in innovative fast-food items.

AI’s Steady Radar: Transforming How Organizations Monitor And Respond

Commercial satellite SAR technology, now widespread in 2026, is revolutionizing how organizations monitor ground changes regardless of weather or time.

Glasspane: One Dataset, Three Views

Glasspane unveils a demo showcasing a single dataset viewed through role-specific perspectives, emphasizing transparency and trust in system monitoring.