🔍 Read the full analysis: AI’s Referee Shortage: A Challenge For Trust And Scale on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A commentary published by ThorstenMeyerAI.com argues that AI is making mathematical, software and contract output faster to produce while expert review remains slow and limited. Its examples point to a potential trust bottleneck, but some figures come from vendors, and the material does not establish that the same pattern applies across every field.
ThorstenMeyerAI.com published an essay this week arguing that AI is making work faster to produce without making it equally easy to trust, leaving human reviewers as a potential bottleneck in mathematics, software and professional workflows. The essay points to OpenAI’s reported production of 722 mathematical manuscripts and studies of AI-assisted coding, while acknowledging that checking results still requires scarce expert judgment.
The essay says OpenAI’s model was given about 4,000 mathematical problems and produced 722 manuscripts grouped into 372 families. Some results were formally checked using Lean, a proof assistant; OpenAI warned that some results without formal verification could have issues. The source contrasts that output with the careful scrutiny by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture. The supplied material does not identify the mathematicians or provide links to the underlying work.
For software, the essay cites figures from Faros AI and LinearB. Faros reported that teams merged 98% more pull requests during high-AI-adoption periods while review time rose 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The essay also cites a peer-reviewed 2026 study finding that 61% of AI-agent pull requests received no human review before being merged or closed. It cautions that several sources sell code-review tools, so their figures warrant care.
In professional services, the essay describes an OpenAI partnership with contract-software company Ironclad. It says GPT-6 Astra averaged 55% of evaluation criteria across 11 tasks, an improvement over its predecessor. The essay’s point is not that the model is unusable, but that a human may still need to identify missed requirements before its work can be relied on.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Human Review Becomes the Constraint
If AI systems produce more drafts, proofs and code than people can check, organisations face a choice between slowing deployment and accepting weaker review. The essay identifies three risks already visible in its cited examples: work merged without review, reviewers deprioritising AI-generated changes, and the producer of a result effectively deciding what deserves attention when outside review is scarce.
This is consequential because verification is not only a technical check. Tests and proof assistants can assess whether work meets specified conditions, but they do not necessarily establish that the conditions match the real need. Contracts still require accountable signatories, and engineering decisions may require professional approval. The essay argues that people able to evaluate output and take responsibility for it could become a capacity limit on AI adoption. That is an interpretation of the cited evidence, not a measured forecast of labour demand or wages.
The source also raises a training concern: junior workers may get fewer opportunities to draft code, proofs or contracts if AI takes on that work. Those tasks have traditionally helped people develop the judgment later needed for review. The material does not quantify whether AI use is already reducing that training, but it identifies a possible tension between short-term productivity and the future supply of experienced reviewers.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Verification Gap
The essay’s central distinction is between producing an answer and establishing that the answer is correct, relevant and safe to use. AI can generate material quickly, while checking may require domain expertise, time and accountability. The contrast is especially clear in mathematics: formal tools can validate a proof against a stated theorem, but human readers may still need to assess whether the theorem is the right one and whether the result matters.
The software figures offer a different kind of evidence. They describe review times, acceptance rates and whether human review occurred, rather than directly measuring the correctness of every change. The results also come from sources with commercial interests in code-review products, a qualification the essay itself raises. Its broader claim—that generation is outpacing verification—draws on a mix of reported operational data and interpretation, not one common measurement across all three fields.
For contracts, the supplied account gives a model evaluation score but no details about the tasks, scoring method or comparison with human performance. The figure therefore signals remaining evaluation criteria; by itself, it does not show how often errors occur in live contracts or what consequences they have.
“Verification abundance, adjudication scarcity.”
— ThorstenMeyerAI.com essay
formal verification software for mathematics
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Available Evidence
The supplied material does not give publication links, full methods or dates for most cited studies, so readers cannot assess their samples and definitions from the essay alone. In particular, the review-time and acceptance figures should not be treated as directly comparable measures unless their methods align. The essay also notes that some sources sell code-review tools, a potential interest readers should keep in mind.
It is not clear how often unreviewed AI-generated work causes errors, how severe any errors are, or whether teams later catch them through testing or other checks. The 55% contract-task score is not accompanied by details of the evaluation criteria. Nor does the material establish whether junior workers are receiving less training as AI adoption grows, or whether organisations are adding new forms of mentorship and review. The essay presents these as risks and open questions rather than settled outcomes.
As an affiliate, we earn on qualifying purchases.
Track Review Practices and Training
The next useful evidence will be transparent follow-up data: how organisations define and record human review, whether review capacity changes as AI adoption rises, and what quality outcomes follow for accepted and unreviewed work. Further detail about the mathematics and contract evaluations—including methods, error rates and independent checks—would also help readers judge how far the examples support the essay’s argument.
Organisations adopting these systems will need to decide how much work requires human approval and how junior staff gain experience in the underlying disciplines. The source offers no policy announcement or specific timetable for those decisions. For now, the development is a reported concern about review capacity and accountability, not proof that human verification has already failed across these fields.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
ThorstenMeyerAI.com published an essay arguing that AI-generated work is growing faster than the human capacity to verify it. It draws examples from mathematics, software development and contract workflows.
Did OpenAI formally verify all 722 mathematical manuscripts?
No. The supplied material says some results were checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state that all 722 manuscripts received formal verification.
What does the cited software data show?
The cited sources report longer waits for review, changes in pull-request acceptance and cases where AI-agent pull requests received no human review. The figures come from different studies and should not be assumed to measure the same thing.
Are human reviewers becoming obsolete?
The essay argues the opposite: expert review may become a constraint as AI output increases. The material does not prove how demand for reviewers or their employment will change.
What remains unknown?
The supplied account does not establish the error rates or real-world harms associated with unreviewed work, nor whether AI is reducing opportunities for junior staff to develop expertise. More details about study methods and follow-up outcomes are needed.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
