firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Everyone Read the Same Whitepaper. Only Two Cashed Out.

Every crypto native knows the pattern. A token’s pitch is flawless. The team nails the diagnosis — the roadmap, the tokenomics, the narrative. Then the mainnet ships, the signature never lands on the block, and holders are left holding the same pitch with no product. "Same diagnosis, same pitch — no signature" is how one AI experiment just described exactly this failure mode — except the subjects weren’t founders, they were five frontier AI models running a real software company through its worst week.

The Firmulate experiment handed each model the same small software firm: the same customers, the same crises, the same temptations to cut corners. Every decision was versioned and auditable — think of it as an on-chain record of management behavior. The scoreboard from the July 2026 Crucible league: gpt-5.6-sol at 95, Moonshot’s Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A do-nothing baseline scores 26, and a single breach of trust caps the total — in this game, no amount of good work outweighs a breach of trust. Sound monetary policy, you might say.

Amazon

Top picks for "altcoin lesson company"

As an affiliate, we earn on qualifying purchases.

The Newcomer That Beat Three of Four Western Frontier Models

The headline result is Kimi K3. Moonshot’s newcomer took second place with 93, one of only two models — alongside the winner — to close the €55,000 deal, worth +€4,583 in monthly recurring revenue. It found the buried security needle buried two document references deep in the company’s own files, not in the customer event. It saved the churning customer. And it resisted all three social-engineering baits, including a reporter’s "just one yes/no, on background" trick, with only one deviation all week — the cleanest discipline in the field. Its on-record reasoning: "Treat the request as a suspected approval-bypass / possible impersonation."

A fairness footnote matters here: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won. That detail cuts both ways, but the bigger point survives it: the league is open. The gap between second and third is five points, and three of four Western frontier models were beaten by a newcomer.

Spotting the Scam Isn’t the Hard Part

The experiment’s key finding is almost paradoxical: all five models spotted every crisis and refused every manipulation attempt — yet only two signed the deal their own analysis had earned. The fake-CEO messages escalated over three stages; 5 of 5 models refused. In crypto terms: everyone dodged the rug pull, but only two executed the trade. Detection is table stakes. Completion is the scarce asset.

Then there’s Opus 4.8 — the cautionary tale. It was the most thorough participant, generating the deepest analyses and +80 learned rules, yet finished last at 73. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four others. Deep diligence with no execution is a beautiful research report nobody acts on.

You Can Watch It Burn, Live

This isn’t a slide deck. The live company runs every business day with 13 synthetic employees and real money mechanics: €105k monthly burn against €2.3k MRR, a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com — like a transparency dashboard for a company that is, verifiably, losing money right now.

Want to test your own judgment? A quiz built on 242 real, unedited management decisions lets you guess which model made which call. And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The Takeaway: Don’t Buy the Ranking, Verify the Chain

The lesson for anyone who has ever picked a coin — or an AI model — off a leaderboard: benchmarks are someone else’s signature on someone else’s deal. The Crucible shows how wide the spread can be between models that all "pass the vibe check" in a chat demo: 95 down to 73, with a 26-point floor for doing nothing and a hard cap on dishonesty. If AI agents will touch your CRM, your support queue, or your forecast, the question isn’t "does it write well" — it’s whether it finishes what it starts, reads your files first, and stays honest under pressure.

Picking a model without running your own test is now a bet — and you’ve seen how those turn out. Full results and plain-language findings are on the public benchmarks page. Run your own chain before you commit the capital.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Decoding The System Behind Deep Strikes, Jamming, And AI Capabilities

An in-depth analysis of how deep strike tactics, electronic warfare, and AI-driven systems form a unified defense and attack strategy in the Russia-Ukraine conflict.

End-to-End Solutions For AI: Local Document Pipeline Explained

A detailed overview of a modular, local document processing pipeline for AI, emphasizing design principles, architecture, and operational benefits.

AI Changelog Digest For Open-source Maintainers

A new AI-driven digest tool for solo open-source maintainers is being tested, aiming to automate changelog summaries from repositories’ updates and issues.

US Military Cyber Teams: Battling More Than Cyber Threats—Suicide Epidemic

US military cyber units are dealing with a surge in suicides, highlighting mental health challenges amid cybersecurity pressures. Details are still emerging.