Why This Benchmark Rewards AI Managers Even At Their Worst
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why This Benchmark Rewards AI Managers Even At Their Worst on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark scores models based on their ability to handle a week of business crises, rewarding partial progress and trustworthiness. The top model scored 95, with a baseline of 26 for minimal effort. This shifts focus toward reliability over perfect performance.

A new benchmark from Firmulate measures how well AI management models perform during a simulated week of business crises, with the top model scoring 95 out of 100. For more details, see the original analysis. The benchmark emphasizes trustworthiness and task completion, rewarding models that do partial work and penalizing breaches of trust. This approach highlights a shift in evaluating AI management effectiveness, especially in real-world business settings.

The benchmark tested four frontier AI models on a small software company’s seven-day crisis scenario, with identical conditions for each. The highest scorer, gpt-5.6-sol, achieved a score of 95, while the baseline, representing minimal effort, scored 26. The scoring system accounts for partial work, such as triaging customer issues and reading documentation, and explicitly penalizes breaches of trust, like unauthorized escalation or manipulation attempts. This approach is discussed in detail in the original analysis. Notably, no model scored a perfect 100, as the designers consider such a score suspicious, indicating potential unmeasured factors or overly idealized performance.

One key finding was that models which thoroughly read their own documentation and refused manipulation attempts performed better, especially in closing deals and maintaining trust. For example, two models successfully closed a €55,000 deal by correctly referencing internal documents, while others failed to do so. The benchmark also included social engineering tests, such as fake CEO messages, which all models refused, indicating robust trust handling. Insights into these tests can be found in the original analysis. However, thoroughness did not always translate into follow-through, as some models hesitated or failed to escalate issues properly, leading to lower scores despite deep rule sets.

At a glance
reportWhen: published July 2026
The developmentA new benchmark developed by Firmulate tests AI management models during a simulated week of business crises, emphasizing trust and task completion over perfect scores.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$81,085▲ 4.3%
Ethereum ETH$2,629▲ 5.5%
Tether USDT$0.9996▲ 0.0%
BNB BNB$761.69▲ 0.8%
XRP XRP$1.42▲ 6.9%
USDC USDC$0.9997▲ 0.0%
Solana SOL$111.9▲ 5.8%
TRON TRX$0.3377▲ 0.5%
Live data · CoinGecko · alternative.me (24h change)
Why This Benchmark Rewards AI Managers Even At Their Worst
AI Management Benchmark · July 2026

Why This Benchmark Rewards AI Managers Even At Their Worst

A new benchmark from Firmulate measures how well AI management models perform during a simulated week of business crises — testing four frontier models on a small software company’s seven-day crisis scenario. Scoring rewards partial progress and trustworthiness over flawless execution.

Top Model Score 95 /100 gpt-5.6-sol, leading the crisis-week simulation
Minimal-Effort Baseline 26 /100 Doing the least possible still has measurable value
Perfect Score Policy 100 = Red Flag Designers treat a flawless 100 as suspicious
Models Tested 4
Crisis Duration 7 Days
Deal Closed €55,000
Fake CEO Attacks 0 Fooled
The Scoring Philosophy

Reliability Over Perfection

The benchmark’s scoring system accounts for partial work — triaging customer issues, reading documentation — and explicitly penalizes breaches of trust like unauthorized escalation or manipulation. No model scored a perfect 100, because the designers consider such a score suspicious.

26
95
Baseline · Minimal Effort Top Score · gpt-5.6-sol 100 · Suspicious / Unmeasured
Key Findings

What Separated Winners From Losers

Trust

Refusing Manipulation Paid Off

Models that read their own documentation thoroughly and refused manipulation attempts performed better, especially in closing deals and maintaining trust.

Execution

Documentation Won the Deal

Two models successfully closed a €55,000 deal by correctly referencing internal documents — while others without that depth failed to close.

Caution

Thoroughness ≠ Follow-Through

Some models hesitated or failed to escalate issues properly, scoring lower despite deep rule sets. Knowledge alone didn’t guarantee action.

Social Engineering

All Models Spotted Fake CEOs

Social engineering tests — including fake CEO messages — were refused by every model, indicating robust trust handling across the frontier.

Integrity

Trust Breaches Heavily Penalized

Unauthorized escalation or manipulation attempts are punished regardless of other performance — trustworthiness is considered paramount.

Anti-Gaming

Every Decision Audited

Suspiciously perfect scores are treated as a red flag; scoring caps ensure models cannot artificially inflate performance without genuine reliability.

Inside The Simulation

Seven Days of Business Crises

Four frontier AI models faced identical conditions in a small software company’s week of escalating crises. The evaluation chain works like this:

1
Setup

Identical company scenario, rules, and documents for all four models.

2
Triage

Customer issues stream in; models must prioritize and partially resolve.

3
Trust Tests

Fake CEO messages and manipulation attempts probe integrity.

4
Deals

Models attempt to close a €55,000 deal using internal documents.

5
Audit

Every decision audited; partial work earns credit, breaches lose it.

Score Landscape

From Baseline To Near-Perfect

The gap between minimal effort and top performance defines the benchmark’s realistic operating range.

gpt-5.6-sol (Top)
95
Frontier Peers
~74
Baseline (Minimal)
26
Reward & Penalty Matrix
Behaviour Score Effect Signal
Triage customer issues✓ RewardedPartial progress counts
Read internal documentation✓ RewardedGrounded decision-making
Refuse manipulation attempts✓ RewardedTrust maintained
Close deal via correct documents✓ RewardedFollow-through
Unauthorized escalation✗ PenalizedBreach of trust
Manipulation attempts✗ PenalizedIntegrity failure
Hesitation / no escalation~ Lower scoreWeak follow-through
Perfect 100 score~ Treated as suspiciousPossible gaming or bias

A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.

— Anonymous Researcher

The benchmark’s scoring system rewards models that complete tasks and uphold trust, even during crises, emphasizing reliability over perfection.

— Thorsten Meyer
Open Questions & Future Directions

What The Week Didn’t Test

Long-Term

Extended Performance

It remains unclear how models perform over longer periods or in complex, less controlled environments. Real-world management involves ongoing, unpredictable challenges.

Systemic

Hidden Vulnerabilities

The focus on partial work and trust might overlook deeper systemic issues or vulnerabilities that could emerge over time in operational deployments.

Next Iterations

Longer, Diverse Tests

Observers expect longer testing periods, more diverse scenarios, pilot programs for companies, and possible regulatory adoption of trust-based metrics.

Key Questions

Frequently Asked

Q1Why reward partial work instead of zero or perfect scores?

The benchmark reflects real-world management, where partial progress and trustworthiness are more valuable than doing nothing or achieving unrealistic perfection. Partial, honest work still provides tangible business value.

Q2What does the baseline score of 26 mean?

It represents minimal effort — triaging emails, reading documents — earning 26 points. Even the least possible work has measurable value, setting a realistic floor for evaluation.

Q3How does trust influence scoring?

Trust breaches such as manipulation attempts or unauthorized escalation are explicitly and heavily penalized, regardless of other performance aspects. Trustworthiness is paramount.

Q4Could models game the system for a higher score?

Every decision is audited and trust is emphasized. Suspiciously perfect scores are a red flag, and scoring caps prevent artificial inflation without genuine reliability.

Q5Will this shape business AI development?

Yes — it pushes developers to prioritize reliability, trustworthiness, and task completion, aligning AI capabilities with actual business needs rather than raw language or problem-solving skills.

Implications for AI Management and Business Trust

This benchmark underscores that in business contexts, the ability of AI managers to complete tasks reliably and maintain trust is more critical than perfect performance. The scoring system rewards partial but honest work, reflecting real-world expectations where trust and consistency matter more than flawless execution. For companies deploying AI in customer support, sales, or operations, this suggests that focus should be on building models that prioritize integrity and task completion over sheer capability. The emphasis on trustworthiness could influence future AI development and evaluation standards, shifting the industry toward more reliable, transparent AI management tools.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Evolution of AI Management Benchmarks

Traditional AI benchmarks primarily measure language fluency, problem-solving, or creative output, often ignoring how models perform in dynamic, real-world management scenarios. The recent development by Firmulate represents a significant shift, introducing a management-focused evaluation that simulates business crises and trust challenges. The idea is to assess how AI models handle incomplete information, manipulative tactics, and the need for follow-through under pressure. This approach responds to industry concerns about deploying AI in operational roles, where partial work and integrity are vital. The benchmark’s design reflects a broader movement toward practical, accountability-oriented AI evaluation, moving beyond theoretical capabilities to assess real-world management skills.

“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”

— an anonymous researcher

Amazon

business crisis management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Long-Term Performance

It remains unclear how these models will perform over extended periods or in more complex, less controlled environments. The benchmark simulates a week of crises, but real-world management involves ongoing, unpredictable challenges. Additionally, the impact of trust breaches on long-term relationships and operational stability has not been fully explored. The scoring system’s emphasis on partial work and trust might overlook deeper systemic issues or vulnerabilities that could emerge over time.

Amazon

trustworthy AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Directions for AI Management Evaluation

Industry observers expect further iterations of the benchmark to include longer testing periods and more diverse scenarios, aiming to better reflect real-world complexity. Developers may refine models to improve follow-through and trust maintenance, driven by the benchmark’s focus. Companies interested in deploying AI management tools might participate in pilot programs to evaluate their own systems against these standards. Additionally, regulatory bodies could adopt similar trust-based metrics to set industry-wide benchmarks for responsible AI use.

Amazon

AI documentation reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why does the benchmark reward partial work instead of zero or perfect scores?

The benchmark aims to reflect real-world management, where partial progress and trustworthiness are more valuable than doing nothing or achieving unrealistic perfection. It recognizes that partial, honest work still provides tangible value in business operations.

What does a score of 26 for the baseline mean?

The baseline represents minimal effort, such as triaging emails or reading documents, and scores 26 points. It shows that even doing the least possible work has measurable value, setting a realistic floor for performance evaluation.

How does trust influence the scoring system?

The benchmark explicitly penalizes breaches of trust, such as manipulation attempts or unauthorized escalation. Trustworthiness is considered paramount, and models that break trust are penalized heavily, regardless of other performance aspects.

Could models cheat or game the system to score higher?

The scoring system is designed to minimize gaming by auditing every decision and emphasizing trust. Suspiciously perfect scores are considered a red flag, and the scoring caps ensure models cannot artificially inflate their performance without genuine reliability.

Will this benchmark influence how AI is developed for business use?

Yes, it encourages developers to prioritize reliability, trustworthiness, and task completion, aligning AI capabilities with actual business needs rather than just language or problem-solving skills.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Radar That Never Blinks: What SAR Actually Does — for Companies, Institutions, and Governments

Explore how Synthetic Aperture Radar (SAR) works, its applications for companies, institutions, and governments, and why it’s transforming remote sensing in 2026.

Technology Operations Signal Monitor: The Future Of Flipper Zero Development

A new signal monitoring tool targets early detection of platform and tooling changes impacting Flipper Zero development, aiding small software teams.

What’s Behind The Buzz? Taco Bell’s Ice Cream Taco And Food Trends

Taco Bell’s new ice cream taco is gaining attention, highlighting emerging food trends and consumer curiosity in innovative fast-food items.

Private AI prompt workspace for sensitive teams

A new private AI prompt workspace tailored for small, regulated teams begins testing, aiming to enhance data control and security in sensitive workflows.