🔍 Read the full analysis: How AI Shifted The Cost From Doing Work To Checking It on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source essay argues that AI has made producing work much cheaper without making it equally cheap to verify. Examples from mathematics, software and contract workflows point to review capacity as a growing constraint, though some figures come from vendors and the broader economic effects remain uncertain.
AI is lowering the cost of producing work in fields including mathematics, software and contract services, but reviewers still need time and expertise to decide whether that output is correct and useful. A recent account from ThorstenMeyerAI.com uses OpenAI’s publication of 722 mathematical manuscripts and software-review data to argue that verification, rather than generation, is becoming a constraint on how much AI work organisations can use.
The source says OpenAI posed about 4,000 mathematical problems and produced 722 manuscripts across 372 families, with an average result taking about three hours of compute. Some results were formally checked using Lean, a proof-assistant system. OpenAI cautioned that some results not formalized in Lean “could have issues.” The count of manuscripts is a reported output figure, not evidence that all results are correct or independently validated.
In software, the article cites analyses from Faros AI and LinearB. Faros reported that teams merged 98% more pull requests during higher-AI-adoption periods, while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organisations found AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. Those are findings reported by the firms, not a single controlled measure of AI’s effects across all software teams.
The source also points to a peer-reviewed 2026 study that found 61% of AI-agent pull requests received no human review before being merged or closed. It says Faros recorded a 31.3% rise in merges with no review during high-adoption periods. The figures describe different datasets and measures; they should not be treated as interchangeable. The article notes that several cited companies sell code-review products, a reason to read their findings with care.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Limits AI Output
If producing drafts, code or research becomes cheaper while expert review remains time-consuming, organisations may be unable to use all the work AI generates. Review capacity becomes a practical limit: teams must decide what to check first, what can be accepted with automated tests and what needs expert scrutiny. The consequences can include longer queues, missed errors or work being approved without adequate examination.
The shift also affects accountability. A proof can be checked against a formal statement, and software can pass a test suite, but people still need to decide whether the statement or tests reflect the actual problem. Contracts, engineering designs and scientific claims carry professional or institutional responsibilities. Human sign-off remains consequential where someone must answer for an error, even when AI helped produce the underlying work.
The source describes a possible “referee premium”: greater demand for people able to assess and take responsibility for AI-assisted work. That is an interpretation, not a measured wage forecast. Whether it translates into higher pay, more hiring or simply heavier workloads will vary by profession and organisation.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Verification Gap
In mathematics, formal verification can establish that a proof follows from its stated assumptions. It does not by itself determine whether the theorem addresses an important question or whether the assumptions are appropriate. The source characterises this as a distinction between checking validity and judging significance. It also refers to a counterexample to an Erdős conjecture that drew careful verification from five leading mathematicians; the supplied material does not provide enough detail to independently establish the full timeline or the reviewers’ conclusions.
Software illustrates the operational problem through pull-request queues and review practices. Automated tests can catch some defects, but their usefulness depends on what they test. The source reports that some reviewers deprioritise AI-generated changes, while other changes pass without human review. These patterns suggest teams are adapting unevenly, not that every AI-written contribution is unreliable.
In professional services, the source describes an OpenAI partnership with contract-software company Ironclad and says GPT-6 Astra averaged 55% of evaluation criteria across 11 tasks, improving on a previous model. That figure indicates performance against the specified evaluation, not readiness to handle every contract independently. The material does not provide the full evaluation methodology or identify the remaining criteria in each task.
““verification abundance, adjudication scarcity.””
— ThorstenMeyerAI.com, describing a recent paper
AI verification tools for software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Figures Cannot Settle
The available material does not establish that AI adoption alone caused the reported changes in review times, acceptance rates or unreviewed merges. The analyses use different organisations, periods and definitions, and some come from vendors with commercial interests in review tools. The numbers are not directly comparable, and the source does not provide enough methodology to assess every limitation.
It is also unclear how much verification work can be automated without shifting risk to another part of the process. Formal systems and tests can check defined properties, but they cannot independently decide whether those properties match a client’s needs, a scientific question or a legal obligation. The pace at which reviewers’ roles, pay and staffing will change is not established by the examples cited.
The longer-term effect on expertise is also uncertain. The source argues that junior staff may lose opportunities to learn by drafting code, proofs and contracts if AI takes over those tasks. That is a credible workforce concern, but the material does not show whether organisations are already seeing a decline in the number of qualified senior reviewers.
As an affiliate, we earn on qualifying purchases.
How Organisations Adapt Review
The immediate test for employers and professional bodies is whether review processes can keep pace with higher output without relying on rubber-stamping or blanket suspicion of AI-assisted work. Teams may need to track review queues, errors and acceptance outcomes separately for different types of work, while making clear who is accountable for final decisions.
For the specific programmes discussed, the next useful evidence would include fuller details on OpenAI’s mathematical results and their verification status, the methodology behind the software datasets, and the criteria used to evaluate GPT-6 Astra. Further disclosures and independent studies will help show whether the reported pattern holds across sectors and over time. Until then, the central development is a reported imbalance: AI can expand production quickly, while trustworthy review still depends on scarce expertise.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The article argues that AI is increasing the volume of work produced in areas such as mathematics and software faster than human review capacity is growing. It presents verification as a developing bottleneck, not a proven universal outcome.
Were all 722 mathematical manuscripts independently verified?
No such conclusion is supported by the supplied material. It says some results were formally checked in Lean and reports OpenAI’s warning that unformalized results “could have issues.”
What did the software analyses report?
Faros AI and LinearB reported changes in pull-request volume, review time, review queues and acceptance rates across their respective datasets. Their measures come from different analyses, so they should not be combined into one estimate of AI’s effect.
Can AI review AI-generated work?
Automated tools can check defined properties, such as whether code passes specified tests or a proof follows formal rules. They may not determine whether the tests or assumptions address the real need, leaving human judgement and accountability important in many settings.
Does the evidence show that reviewers will earn more?
No. The source suggests that skilled reviewers could become more valuable as AI output grows, but it gives no wage forecast or direct evidence that pay is increasing. Effects on hiring, compensation and workload remain uncertain.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
