How To Think About AI’s Growing Verification Burden
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: How To Think About AI’s Growing Verification Burden on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get hardware and tech essentials delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A source article argues that AI is making it cheaper to produce research, code and professional work while human review remains limited. Examples from mathematics, software and contract workflows point to a widening gap, though several cited software metrics come from vendors and should be read with care.

AI-generated work is expanding faster than review capacity, according to an analysis published this week by ThorstenMeyerAI.com. The article uses OpenAI’s release of 722 mathematical manuscripts as its lead example, arguing that producing results is becoming cheaper while verifying their correctness and usefulness remains a slower, expert-dependent process.

The source says OpenAI’s model was given about 4,000 mathematical problems and produced 722 manuscripts grouped into 372 families, with an average result taking about three hours of compute. Some results were formally checked using Lean, a proof-assistant system. OpenAI cautioned, according to the article, that some results without formal verification “could have issues.” The source contrasts that output with the careful review by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an Erdős conjecture.

In software, the article cites several measures of review workload. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time increased 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. The article notes that some cited data comes from companies selling code-review tools.

A third example concerns professional work. The source says OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on contracting workflows. Across 11 tasks, the model met an average of 55% of evaluation criteria, which the article describes as an improvement over its predecessor. The remaining criteria illustrate the need for review, but the source does not provide the evaluation methodology or identify each missed criterion.

At a glance
analysisWhen: Published this week, according to the s…
The developmentA commentary published by ThorstenMeyerAI.com argues that AI’s rising output is increasing the burden on people who must check and take responsibility for it.
Crypto market snapshot
Fear & Greed Index
64/100 — Greed
Bitcoin BTC$82,943▼ 1.2%
Ethereum ETH$2,571▼ 1.4%
Tether USDT$0.9994▼ 0.0%
BNB BNB$770.53▲ 0.6%
XRP XRP$1.42▼ 2.9%
USDC USDC$0.9996▼ 0.0%
Solana SOL$115.93▼ 1.9%
TRON TRX$0.3357▲ 1.0%
Live data · CoinGecko · alternative.me (24h change)
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Becomes a Constraint

The analysis frames verification as a potential operational bottleneck: organisations may generate more code, research and contract drafts than qualified people can reliably assess. If review lags, work can be delayed, accepted with insufficient scrutiny or filtered mainly by the system or team that produced it. These are risks described by the source, not proof that every AI-assisted workflow is failing.

The article also points to a workforce concern. Junior professionals often build expertise by doing the very tasks that AI tools can now accelerate or replace: writing code, drafting documents and developing arguments. If training shifts away from that work without creating another route to practice, organisations could have fewer experienced reviewers just as demand for their judgment grows. The source calls this a possible “referee premium” for people able to assess work and take responsibility for approving it.

That argument matters beyond productivity measures. Review is not only error detection; it also helps determine whether the work answers the right question and whether it is safe or suitable to use. In fields such as contracts, engineering and research, final responsibility rests with people or institutions, not with a model.

Amazon

AI verification software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, Similar Review Pressures

The article links mathematics, software and contracting through one distinction: producing an output and establishing that it is trustworthy are separate tasks. Formal tools can check some properties. A proof assistant can test whether a proof follows from stated assumptions, and software tests can check specified behaviours. But those checks do not establish that the theorem, requirements or tests address the real problem.

The source describes this as “verification abundance, adjudication scarcity”: more candidate work is available for checking, while expert judgment remains limited. It also reports that reviewers may find AI-written code more taxing to assess because the output does not reveal the author’s reasoning or where mistakes are likely. That is the source’s explanation, rather than a universal finding established by the examples alone.

The evidence cited is mixed in type and scope. It includes company-reported metrics, a peer-reviewed study, an OpenAI project and an Ironclad partnership. The figures therefore do not form a single controlled comparison across industries, and the software data has a possible commercial interest attached to some sources.

Amazon

code review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the Available Metrics Cannot Show

The cited figures do not establish how much AI-generated work contains consequential errors, or whether faster production leads to better or worse outcomes overall. The source does not provide detailed methods for every company metric, and it acknowledges that some data comes from vendors with a commercial interest in review tooling. The figures should not be treated as directly comparable measures of one industry-wide trend.

It is also unclear how much of the verification burden can be reduced through automated checking, improved model reliability or changes to workflow design. The source argues that human review will remain necessary for intent, relevance and accountability, but it does not quantify how many expert hours are required or how those needs differ across fields.

The article’s concern about future reviewer shortages is a projection, not a confirmed workforce outcome. It does not establish whether junior staff will lose training opportunities, or whether organisations can preserve those opportunities through supervised AI use and other forms of practice.

Amazon

formal proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Organisations Adapt Review

The source does not identify a scheduled policy change or next milestone. The immediate question for organisations is how they will match review capacity to AI-assisted output: what work requires human sign-off, which checks can be automated, and who remains accountable for decisions. The cited examples make those choices relevant now, but they do not show that one approach has been established as best practice.

Further evidence will be needed to test whether the reported review delays and acceptance rates persist across sectors and over time. Useful comparisons would distinguish AI-assisted work by task and risk, disclose how review outcomes are measured, and track whether teams maintain training routes for junior staff. Until then, the article’s central claim is a warning about a possible imbalance: generation is scaling quickly, while trustworthy adjudication still depends heavily on experienced people.

Amazon

AI review management platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main argument of the analysis?

It argues that AI can increase the volume of work more quickly than people can verify it, leaving review and accountability as potential bottlenecks.

Did OpenAI formally verify all 722 mathematical manuscripts?

No. The source says some results were checked in Lean and reports OpenAI’s warning that some unformalized results “could have issues.” It does not say all manuscripts received formal verification.

How reliable are the software review figures?

The source cites Faros AI, LinearB and a peer-reviewed 2026 study, but notes that some sources sell code-review tools. The figures have different scopes and methods, so they should be read as separate reported findings rather than a single definitive measure.

Why does the analysis say human review remains necessary?

It says automated checks can test whether work meets specified rules without necessarily confirming that the rules address the right need. It also points to professional and legal accountability, which remains assigned to people or institutions.

Does the article show that AI is eliminating reviewer jobs?

No. It raises the possibility that AI could weaken some routes through which junior workers gain expertise, while increasing demand for experienced reviewers. It does not establish that this outcome has already occurred.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Gulf: Own the Capital

Gulf states are investing heavily in AI infrastructure to own the next economy, using oil wealth for strategic capital deployment amid shifting global dynamics.

The stake. Why the answer to automation is broad-based ownership, not a bigger transfer.

Thorsten Meyer argues that expanding ownership of capital, not increasing transfers, is the market-friendly way to address automation’s economic shifts.

After the Paycheck: The Book I Wrote Because Nobody Else Would Tell the Truth About AI and Your Income

Author Thorsten Meyer releases ‘After the Paycheck,’ analyzing AI’s impact on jobs, ownership, and economic security, emphasizing ownership over automation.

The Quiet Audit: 55–75% of Your Week Is on Thin Ice. Here’s Which Part.

Research shows 55–75% of knowledge workers’ time is on thin ice, mostly involving theatre, commodity, or on-the-line tasks. Here’s what you need to know.