🔍 Read the full analysis: Why This Benchmark Rewards AI Managers Even At Their Worst on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark scores models based on their ability to handle a week of business crises, rewarding partial progress and trustworthiness. The top model scored 95, with a baseline of 26 for minimal effort. This shifts focus toward reliability over perfect performance.
A new benchmark from Firmulate measures how well AI management models perform during a simulated week of business crises, with the top model scoring 95 out of 100. For more details, see the original analysis. The benchmark emphasizes trustworthiness and task completion, rewarding models that do partial work and penalizing breaches of trust. This approach highlights a shift in evaluating AI management effectiveness, especially in real-world business settings.
The benchmark tested four frontier AI models on a small software company’s seven-day crisis scenario, with identical conditions for each. The highest scorer, gpt-5.6-sol, achieved a score of 95, while the baseline, representing minimal effort, scored 26. The scoring system accounts for partial work, such as triaging customer issues and reading documentation, and explicitly penalizes breaches of trust, like unauthorized escalation or manipulation attempts. This approach is discussed in detail in the original analysis. Notably, no model scored a perfect 100, as the designers consider such a score suspicious, indicating potential unmeasured factors or overly idealized performance.
One key finding was that models which thoroughly read their own documentation and refused manipulation attempts performed better, especially in closing deals and maintaining trust. For example, two models successfully closed a €55,000 deal by correctly referencing internal documents, while others failed to do so. The benchmark also included social engineering tests, such as fake CEO messages, which all models refused, indicating robust trust handling. Insights into these tests can be found in the original analysis. However, thoroughness did not always translate into follow-through, as some models hesitated or failed to escalate issues properly, leading to lower scores despite deep rule sets.
Why This Benchmark Rewards AI Managers Even At Their Worst
A new benchmark from Firmulate measures how well AI management models perform during a simulated week of business crises — testing four frontier models on a small software company’s seven-day crisis scenario. Scoring rewards partial progress and trustworthiness over flawless execution.
Reliability Over Perfection
The benchmark’s scoring system accounts for partial work — triaging customer issues, reading documentation — and explicitly penalizes breaches of trust like unauthorized escalation or manipulation. No model scored a perfect 100, because the designers consider such a score suspicious.
What Separated Winners From Losers
Refusing Manipulation Paid Off
Models that read their own documentation thoroughly and refused manipulation attempts performed better, especially in closing deals and maintaining trust.
Documentation Won the Deal
Two models successfully closed a €55,000 deal by correctly referencing internal documents — while others without that depth failed to close.
Thoroughness ≠ Follow-Through
Some models hesitated or failed to escalate issues properly, scoring lower despite deep rule sets. Knowledge alone didn’t guarantee action.
All Models Spotted Fake CEOs
Social engineering tests — including fake CEO messages — were refused by every model, indicating robust trust handling across the frontier.
Trust Breaches Heavily Penalized
Unauthorized escalation or manipulation attempts are punished regardless of other performance — trustworthiness is considered paramount.
Every Decision Audited
Suspiciously perfect scores are treated as a red flag; scoring caps ensure models cannot artificially inflate performance without genuine reliability.
Seven Days of Business Crises
Four frontier AI models faced identical conditions in a small software company’s week of escalating crises. The evaluation chain works like this:
Identical company scenario, rules, and documents for all four models.
Customer issues stream in; models must prioritize and partially resolve.
Fake CEO messages and manipulation attempts probe integrity.
Models attempt to close a €55,000 deal using internal documents.
Every decision audited; partial work earns credit, breaches lose it.
From Baseline To Near-Perfect
The gap between minimal effort and top performance defines the benchmark’s realistic operating range.
Reward & Penalty Matrix| Behaviour | Score Effect | Signal |
|---|---|---|
| Triage customer issues | ✓ Rewarded | Partial progress counts |
| Read internal documentation | ✓ Rewarded | Grounded decision-making |
| Refuse manipulation attempts | ✓ Rewarded | Trust maintained |
| Close deal via correct documents | ✓ Rewarded | Follow-through |
| Unauthorized escalation | ✗ Penalized | Breach of trust |
| Manipulation attempts | ✗ Penalized | Integrity failure |
| Hesitation / no escalation | ~ Lower score | Weak follow-through |
| Perfect 100 score | ~ Treated as suspicious | Possible gaming or bias |
A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.
— Anonymous ResearcherThe benchmark’s scoring system rewards models that complete tasks and uphold trust, even during crises, emphasizing reliability over perfection.
— Thorsten MeyerWhat The Week Didn’t Test
Extended Performance
It remains unclear how models perform over longer periods or in complex, less controlled environments. Real-world management involves ongoing, unpredictable challenges.
Hidden Vulnerabilities
The focus on partial work and trust might overlook deeper systemic issues or vulnerabilities that could emerge over time in operational deployments.
Longer, Diverse Tests
Observers expect longer testing periods, more diverse scenarios, pilot programs for companies, and possible regulatory adoption of trust-based metrics.
Frequently Asked
Q1Why reward partial work instead of zero or perfect scores?
The benchmark reflects real-world management, where partial progress and trustworthiness are more valuable than doing nothing or achieving unrealistic perfection. Partial, honest work still provides tangible business value.
Q2What does the baseline score of 26 mean?
It represents minimal effort — triaging emails, reading documents — earning 26 points. Even the least possible work has measurable value, setting a realistic floor for evaluation.
Q3How does trust influence scoring?
Trust breaches such as manipulation attempts or unauthorized escalation are explicitly and heavily penalized, regardless of other performance aspects. Trustworthiness is paramount.
Q4Could models game the system for a higher score?
Every decision is audited and trust is emphasized. Suspiciously perfect scores are a red flag, and scoring caps prevent artificial inflation without genuine reliability.
Q5Will this shape business AI development?
Yes — it pushes developers to prioritize reliability, trustworthiness, and task completion, aligning AI capabilities with actual business needs rather than raw language or problem-solving skills.
Implications for AI Management and Business Trust
This benchmark underscores that in business contexts, the ability of AI managers to complete tasks reliably and maintain trust is more critical than perfect performance. The scoring system rewards partial but honest work, reflecting real-world expectations where trust and consistency matter more than flawless execution. For companies deploying AI in customer support, sales, or operations, this suggests that focus should be on building models that prioritize integrity and task completion over sheer capability. The emphasis on trustworthiness could influence future AI development and evaluation standards, shifting the industry toward more reliable, transparent AI management tools.
As an affiliate, we earn on qualifying purchases.
The Evolution of AI Management Benchmarks
Traditional AI benchmarks primarily measure language fluency, problem-solving, or creative output, often ignoring how models perform in dynamic, real-world management scenarios. The recent development by Firmulate represents a significant shift, introducing a management-focused evaluation that simulates business crises and trust challenges. The idea is to assess how AI models handle incomplete information, manipulative tactics, and the need for follow-through under pressure. This approach responds to industry concerns about deploying AI in operational roles, where partial work and integrity are vital. The benchmark’s design reflects a broader movement toward practical, accountability-oriented AI evaluation, moving beyond theoretical capabilities to assess real-world management skills.
“A manager who does something useful is not the same as one who does nothing, and pretending otherwise would make the benchmark dishonest.”
— an anonymous researcher
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Long-Term Performance
It remains unclear how these models will perform over extended periods or in more complex, less controlled environments. The benchmark simulates a week of crises, but real-world management involves ongoing, unpredictable challenges. Additionally, the impact of trust breaches on long-term relationships and operational stability has not been fully explored. The scoring system’s emphasis on partial work and trust might overlook deeper systemic issues or vulnerabilities that could emerge over time.
trustworthy AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Future Directions for AI Management Evaluation
Industry observers expect further iterations of the benchmark to include longer testing periods and more diverse scenarios, aiming to better reflect real-world complexity. Developers may refine models to improve follow-through and trust maintenance, driven by the benchmark’s focus. Companies interested in deploying AI management tools might participate in pilot programs to evaluate their own systems against these standards. Additionally, regulatory bodies could adopt similar trust-based metrics to set industry-wide benchmarks for responsible AI use.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why does the benchmark reward partial work instead of zero or perfect scores?
The benchmark aims to reflect real-world management, where partial progress and trustworthiness are more valuable than doing nothing or achieving unrealistic perfection. It recognizes that partial, honest work still provides tangible value in business operations.
What does a score of 26 for the baseline mean?
The baseline represents minimal effort, such as triaging emails or reading documents, and scores 26 points. It shows that even doing the least possible work has measurable value, setting a realistic floor for performance evaluation.
How does trust influence the scoring system?
The benchmark explicitly penalizes breaches of trust, such as manipulation attempts or unauthorized escalation. Trustworthiness is considered paramount, and models that break trust are penalized heavily, regardless of other performance aspects.
Could models cheat or game the system to score higher?
The scoring system is designed to minimize gaming by auditing every decision and emphasizing trust. Suspiciously perfect scores are considered a red flag, and the scoring caps ensure models cannot artificially inflate their performance without genuine reliability.
Will this benchmark influence how AI is developed for business use?
Yes, it encourages developers to prioritize reliability, trustworthiness, and task completion, aligning AI capabilities with actual business needs rather than just language or problem-solving skills.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
