🔍 Read the full analysis: AI Makes Creating Easier. Is Checking Becoming The Hard Part? on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
AI systems are producing mathematical manuscripts, software changes and contract work faster, while review remains a human bottleneck. The source material cites studies and company data suggesting review is taking longer or being skipped, but some figures come from vendors that sell review tools and should be read with care.
AI-generated work is arriving faster than people can verify it, according to a report drawing on new mathematical output, software-development data and a contract-workflow evaluation. The examples point to a growing gap between the low cost of generating drafts or code and the expert time required to establish that the results are correct and fit for use.
The report says OpenAI published 722 mathematical manuscripts, grouped into 372 families, after its model was given about 4,000 problems. It estimates that the average result took about three hours of compute to produce. OpenAI said some results were formally checked using Lean, a proof assistant, and cautioned that some unformalized results “could have issues.” The source does not provide a full breakdown of which manuscripts received which forms of review.
In software, the report cites figures from several sources. Faros AI reported that teams in periods of high AI adoption merged 98% more pull requests, while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source also cites a 2026 peer-reviewed study that found 61% of AI-agent pull requests received no human review before they were merged or closed.
A separate example comes from professional services. The report says OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on contracting workflows. In an evaluation across 11 tasks, Astra met 55% of the criteria on average, according to the source. That result indicates improvement over a prior model, but it also leaves a substantial share of evaluation criteria unmet. The material does not give the evaluation methodology or identify the previous model’s score.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Is Becoming the Capacity Limit
If organizations can create more work than they can responsibly inspect, the constraint shifts from production to approval. Review time can determine how much AI output is usable, whether a software change is safe to merge, or whether a contract draft is ready for a person to sign. Faster generation alone does not show that the work is accurate or fit for its intended purpose.
The source describes several possible responses already visible in the cited data: some changes may be merged without review, reviewers may deprioritize AI-generated work, or the organization producing the work may decide which items merit outside scrutiny. Each approach has trade-offs. Skipping review can leave errors undetected; treating all AI output as suspect can slow useful work; and relying on the producer’s own selection leaves less independent scrutiny.
The report also raises a workforce concern: review expertise is built through practice. If junior staff spend less time drafting, coding or proving work themselves, they may get fewer opportunities to develop the judgement expected of senior reviewers. Whether AI adoption will produce that effect at scale is not established by the examples provided, but it is a relevant question for employers training the next generation of specialists.
AI review tools for software development
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three Fields Show a Similar Gap
The examples span mathematics, software and contract work, but the underlying distinction is similar: automated systems can produce a candidate result, while checking whether it answers the right question may still require people. In mathematics, formal tools can verify that a proof follows from a stated theorem. That does not, by itself, establish that the theorem is the intended one or that the result is important. In software, tests can check the behavior they cover, but cannot guarantee that the tests represent every real requirement.
The source says five leading mathematicians carefully verified an earlier result from the same program: a counterexample to an old Erdős conjecture. It presents that response as a contrast with the volume of newly generated manuscripts. The specific review process and the exact scope of the five mathematicians’ work are not detailed in the supplied material.
Some figures come from companies that sell software review products, including Faros AI and LinearB. That commercial interest is a reason to examine their methods and definitions, not by itself a reason to dismiss their findings. The report says the studies point in a consistent direction, but it does not provide links, full methods or enough detail to independently compare how each source defined review time, acceptance or AI-generated work.
How Much Review Is Being Skipped?
The cited evidence does not establish a single, comparable measure of AI verification across all three fields. The source does not identify the mathematical manuscripts that were formally checked, give full details of the contract evaluation, or provide study methods for every software statistic. The numbers should be understood within their stated samples and definitions, not as universal rates for all organizations or AI tools.
It is also unclear whether increased review times are caused by AI-generated work itself, higher overall submission volume, changes in team practices, or a combination of factors. The report says some reviewers deliberately deprioritize AI-generated changes, but the cited material does not establish how common that practice is beyond the surveyed data. Nor does it show whether organizations are adding review staff or changing processes to address the backlog.
Finally, the proposed “referee premium”—greater value for people able to approve work and take responsibility for it—is an interpretation, not a measured labor-market outcome in the material provided. Evidence about hiring, compensation and training would be needed to determine whether that pattern is already occurring.
Measure Quality Alongside Output
The next test for organizations adopting AI is whether they can track the quality and review status of output, not only how much is produced or how quickly it moves through a workflow. That means making clear what was checked, by whom, against which requirements, and what remains uncertain before work is accepted or used.
For the math results, readers will need clearer information about which manuscripts were formally verified and what independent review they received. For software, further studies with transparent definitions and methods could help distinguish the effect of AI-generated changes from shifts in submission volume or team organization. In contract workflows, published evaluation details would help clarify what the 55% score measures and how performance compares with human review or earlier systems.
The source material does not identify a specific next milestone or timetable. For now, the central issue is whether human review capacity, training and accountability can keep pace with the growing volume of AI-assisted work.
Key Questions
What is the main development described?
The report says AI systems are producing large volumes of mathematical, software and contract-related work, while people remain responsible for checking whether it is correct and suitable for use.
Did OpenAI verify all 722 mathematical manuscripts?
The source does not say that all 722 were formally verified. It says some results were checked in Lean and quotes OpenAI warning that some unformalized results “could have issues.”
What did the cited software data find?
The report cites analyses that found more pull requests in high-AI-adoption periods, longer waits for review of AI-generated changes and some changes receiving no human review. The figures refer to particular datasets and definitions, and some sources sell review tools.
Does the report prove that AI is reducing the supply of expert reviewers?
No. It raises the possibility that less hands-on work for junior staff could weaken the experience through which reviewers develop judgement. The supplied evidence does not establish a measured, economy-wide decline in reviewer supply.
What remains unknown about the contract evaluation?
The source reports that GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, but does not provide the evaluation method, the previous model’s score or details about how the criteria were weighted.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
