🔍 Read the full analysis: How To Think About AI’s Growing Verification Burden on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source article argues that AI is making it cheaper to produce research, code and professional work while human review remains limited. Examples from mathematics, software and contract workflows point to a widening gap, though several cited software metrics come from vendors and should be read with care.
AI-generated work is expanding faster than review capacity, according to an analysis published this week by ThorstenMeyerAI.com. The article uses OpenAI’s release of 722 mathematical manuscripts as its lead example, arguing that producing results is becoming cheaper while verifying their correctness and usefulness remains a slower, expert-dependent process.
The source says OpenAI’s model was given about 4,000 mathematical problems and produced 722 manuscripts grouped into 372 families, with an average result taking about three hours of compute. Some results were formally checked using Lean, a proof-assistant system. OpenAI cautioned, according to the article, that some results without formal verification “could have issues.” The source contrasts that output with the careful review by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an Erdős conjecture.
In software, the article cites several measures of review workload. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time increased 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before being merged or closed. The article notes that some cited data comes from companies selling code-review tools.
A third example concerns professional work. The source says OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on contracting workflows. Across 11 tasks, the model met an average of 55% of evaluation criteria, which the article describes as an improvement over its predecessor. The remaining criteria illustrate the need for review, but the source does not provide the evaluation methodology or identify each missed criterion.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Becomes a Constraint
The analysis frames verification as a potential operational bottleneck: organisations may generate more code, research and contract drafts than qualified people can reliably assess. If review lags, work can be delayed, accepted with insufficient scrutiny or filtered mainly by the system or team that produced it. These are risks described by the source, not proof that every AI-assisted workflow is failing.
The article also points to a workforce concern. Junior professionals often build expertise by doing the very tasks that AI tools can now accelerate or replace: writing code, drafting documents and developing arguments. If training shifts away from that work without creating another route to practice, organisations could have fewer experienced reviewers just as demand for their judgment grows. The source calls this a possible “referee premium” for people able to assess work and take responsibility for approving it.
That argument matters beyond productivity measures. Review is not only error detection; it also helps determine whether the work answers the right question and whether it is safe or suitable to use. In fields such as contracts, engineering and research, final responsibility rests with people or institutions, not with a model.
As an affiliate, we earn on qualifying purchases.
Three Fields, Similar Review Pressures
The article links mathematics, software and contracting through one distinction: producing an output and establishing that it is trustworthy are separate tasks. Formal tools can check some properties. A proof assistant can test whether a proof follows from stated assumptions, and software tests can check specified behaviours. But those checks do not establish that the theorem, requirements or tests address the real problem.
The source describes this as “verification abundance, adjudication scarcity”: more candidate work is available for checking, while expert judgment remains limited. It also reports that reviewers may find AI-written code more taxing to assess because the output does not reveal the author’s reasoning or where mistakes are likely. That is the source’s explanation, rather than a universal finding established by the examples alone.
The evidence cited is mixed in type and scope. It includes company-reported metrics, a peer-reviewed study, an OpenAI project and an Ironclad partnership. The figures therefore do not form a single controlled comparison across industries, and the software data has a possible commercial interest attached to some sources.
As an affiliate, we earn on qualifying purchases.
What the Available Metrics Cannot Show
The cited figures do not establish how much AI-generated work contains consequential errors, or whether faster production leads to better or worse outcomes overall. The source does not provide detailed methods for every company metric, and it acknowledges that some data comes from vendors with a commercial interest in review tooling. The figures should not be treated as directly comparable measures of one industry-wide trend.
It is also unclear how much of the verification burden can be reduced through automated checking, improved model reliability or changes to workflow design. The source argues that human review will remain necessary for intent, relevance and accountability, but it does not quantify how many expert hours are required or how those needs differ across fields.
The article’s concern about future reviewer shortages is a projection, not a confirmed workforce outcome. It does not establish whether junior staff will lose training opportunities, or whether organisations can preserve those opportunities through supervised AI use and other forms of practice.
As an affiliate, we earn on qualifying purchases.
How Organisations Adapt Review
The source does not identify a scheduled policy change or next milestone. The immediate question for organisations is how they will match review capacity to AI-assisted output: what work requires human sign-off, which checks can be automated, and who remains accountable for decisions. The cited examples make those choices relevant now, but they do not show that one approach has been established as best practice.
Further evidence will be needed to test whether the reported review delays and acceptance rates persist across sectors and over time. Useful comparisons would distinguish AI-assisted work by task and risk, disclose how review outcomes are measured, and track whether teams maintain training routes for junior staff. Until then, the article’s central claim is a warning about a possible imbalance: generation is scaling quickly, while trustworthy adjudication still depends heavily on experienced people.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main argument of the analysis?
It argues that AI can increase the volume of work more quickly than people can verify it, leaving review and accountability as potential bottlenecks.
Did OpenAI formally verify all 722 mathematical manuscripts?
No. The source says some results were checked in Lean and reports OpenAI’s warning that some unformalized results “could have issues.” It does not say all manuscripts received formal verification.
How reliable are the software review figures?
The source cites Faros AI, LinearB and a peer-reviewed 2026 study, but notes that some sources sell code-review tools. The figures have different scopes and methods, so they should be read as separate reported findings rather than a single definitive measure.
Why does the analysis say human review remains necessary?
It says automated checks can test whether work meets specified rules without necessarily confirming that the rules address the right need. It also points to professional and legal accountability, which remains assigned to people or institutions.
Does the article show that AI is eliminating reviewer jobs?
No. It raises the possibility that AI could weaken some routes through which junior workers gain expertise, while increasing demand for experienced reviewers. It does not establish that this outcome has already occurred.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
