AI’s Referee Shortage: A Challenge For Trust And Scale
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI’s Referee Shortage: A Challenge For Trust And Scale on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get hardware and tech essentials delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A commentary published by ThorstenMeyerAI.com argues that AI is making mathematical, software and contract output faster to produce while expert review remains slow and limited. Its examples point to a potential trust bottleneck, but some figures come from vendors, and the material does not establish that the same pattern applies across every field.

ThorstenMeyerAI.com published an essay this week arguing that AI is making work faster to produce without making it equally easy to trust, leaving human reviewers as a potential bottleneck in mathematics, software and professional workflows. The essay points to OpenAI’s reported production of 722 mathematical manuscripts and studies of AI-assisted coding, while acknowledging that checking results still requires scarce expert judgment.

The essay says OpenAI’s model was given about 4,000 mathematical problems and produced 722 manuscripts grouped into 372 families. Some results were formally checked using Lean, a proof assistant; OpenAI warned that some results without formal verification could have issues. The source contrasts that output with the careful scrutiny by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture. The supplied material does not identify the mathematicians or provide links to the underlying work.

For software, the essay cites figures from Faros AI and LinearB. Faros reported that teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The essay also cites a peer-reviewed 2026 study finding that 61% of AI-agent pull requests received no human review before being merged or closed. It cautions that several sources sell code-review tools, so their figures warrant care.

In professional services, the essay describes an OpenAI partnership with contract-software company Ironclad. It says GPT-6 Astra averaged 55% of evaluation criteria across 11 tasks, an improvement over its predecessor. The essay’s point is not that the model is unusable, but that a human may still need to identify missed requirements before its work can be relied on.

At a glance
reportWhen: Published this week, according to the s…
The developmentThorstenMeyerAI.com published an essay arguing that AI-generated work is increasing faster than the human capacity to verify it.
Crypto market snapshot
Fear & Greed Index
64/100 — Greed
Bitcoin BTC$82,846▼ 1.6%
Ethereum ETH$2,567▼ 1.7%
Tether USDT$0.9994▼ 0.0%
BNB BNB$769.91▲ 0.3%
XRP XRP$1.42▼ 3.2%
USDC USDC$0.9996▼ 0.0%
Solana SOL$115.89▼ 2.0%
TRON TRX$0.3356▲ 0.9%
Live data · CoinGecko · alternative.me (24h change)
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Human Review Becomes the Constraint

If AI systems produce more drafts, proofs and code than people can check, organisations face a choice between slowing deployment and accepting weaker review. The essay identifies three risks already visible in its cited examples: work merged without review, reviewers deprioritising AI-generated changes, and the producer of a result effectively deciding what deserves attention when outside review is scarce.

This is consequential because verification is not only a technical check. Tests and proof assistants can assess whether work meets specified conditions, but they do not necessarily establish that the conditions match the real need. Contracts still require accountable signatories, and engineering decisions may require professional approval. The essay argues that people able to evaluate output and take responsibility for it could become a capacity limit on AI adoption. That is an interpretation of the cited evidence, not a measured forecast of labour demand or wages.

The source also raises a training concern: junior workers may get fewer opportunities to draft code, proofs or contracts if AI takes on that work. Those tasks have traditionally helped people develop the judgment later needed for review. The material does not quantify whether AI use is already reducing that training, but it identifies a possible tension between short-term productivity and the future supply of experienced reviewers.

Amazon

AI code review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Verification Gap

The essay’s central distinction is between producing an answer and establishing that the answer is correct, relevant and safe to use. AI can generate material quickly, while checking may require domain expertise, time and accountability. The contrast is especially clear in mathematics: formal tools can validate a proof against a stated theorem, but human readers may still need to assess whether the theorem is the right one and whether the result matters.

The software figures offer a different kind of evidence. They describe review times, acceptance rates and whether human review occurred, rather than directly measuring the correctness of every change. The results also come from sources with commercial interests in code-review products, a qualification the essay itself raises. Its broader claim—that generation is outpacing verification—draws on a mix of reported operational data and interpretation, not one common measurement across all three fields.

For contracts, the supplied account gives a model evaluation score but no details about the tasks, scoring method or comparison with human performance. The figure therefore signals remaining evaluation criteria; by itself, it does not show how often errors occur in live contracts or what consequences they have.

“Verification abundance, adjudication scarcity.”

— ThorstenMeyerAI.com essay

Amazon

formal verification software for mathematics

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Available Evidence

The supplied material does not give publication links, full methods or dates for most cited studies, so readers cannot assess their samples and definitions from the essay alone. In particular, the review-time and acceptance figures should not be treated as directly comparable measures unless their methods align. The essay also notes that some sources sell code-review tools, a potential interest readers should keep in mind.

It is not clear how often unreviewed AI-generated work causes errors, how severe any errors are, or whether teams later catch them through testing or other checks. The 55% contract-task score is not accompanied by details of the evaluation criteria. Nor does the material establish whether junior workers are receiving less training as AI adoption grows, or whether organisations are adding new forms of mentorship and review. The essay presents these as risks and open questions rather than settled outcomes.

Amazon

AI contract review software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review Practices and Training

The next useful evidence will be transparent follow-up data: how organisations define and record human review, whether review capacity changes as AI adoption rises, and what quality outcomes follow for accepted and unreviewed work. Further detail about the mathematics and contract evaluations—including methods, error rates and independent checks—would also help readers judge how far the examples support the essay’s argument.

Organisations adopting these systems will need to decide how much work requires human approval and how junior staff gain experience in the underlying disciplines. The source offers no policy announcement or specific timetable for those decisions. For now, the development is a reported concern about review capacity and accountability, not proof that human verification has already failed across these fields.

Amazon

proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

ThorstenMeyerAI.com published an essay arguing that AI-generated work is growing faster than the human capacity to verify it. It draws examples from mathematics, software development and contract workflows.

Did OpenAI formally verify all 722 mathematical manuscripts?

No. The supplied material says some results were checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state that all 722 manuscripts received formal verification.

What does the cited software data show?

The cited sources report longer waits for review, changes in pull-request acceptance and cases where AI-agent pull requests received no human review. The figures come from different studies and should not be assumed to measure the same thing.

Are human reviewers becoming obsolete?

The essay argues the opposite: expert review may become a constraint as AI output increases. The material does not prove how demand for reviewers or their employment will change.

What remains unknown?

The supplied account does not establish the error rates or real-world harms associated with unreviewed work, nor whether AI is reducing opportunities for junior staff to develop expertise. More details about study methods and follow-up outcomes are needed.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Docking Stations Simplify Multi-Monitor Crypto Setups

Meta description: Maximizing your crypto workflow is easier with docking stations, but discover how they truly transform your multi-monitor setup and why they matter.

Why Mining Power Supplies Deserve More Attention Than Most Buyers Give Them

Discover why mining power supplies demand your full attention to ensure optimal performance and avoid costly issues down the line.

9 Best Camera Lenses In 2026

Discover the nine best camera lenses of 2026, covering versatile zooms, primes, telephotos, and more. Find the perfect lens for your photography style.

Future Focus: 10 AI Trends That Will Take Off In 2026

A forecast of the 10 key AI trends expected to emerge in 2026, based on industry insights and expert analysis, highlighting potential impacts and uncertainties.