AI Makes Creating Easier. Is Checking Becoming The Hard Part?
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: AI Makes Creating Easier. Is Checking Becoming The Hard Part? on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on office and shipping supplies

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

AI systems are producing mathematical manuscripts, software changes and contract work faster, while review remains a human bottleneck. The source material cites studies and company data suggesting review is taking longer or being skipped, but some figures come from vendors that sell review tools and should be read with care.

AI-generated work is arriving faster than people can verify it, according to a report drawing on new mathematical output, software-development data and a contract-workflow evaluation. The examples point to a growing gap between the low cost of generating drafts or code and the expert time required to establish that the results are correct and fit for use.

The report says OpenAI published 722 mathematical manuscripts, grouped into 372 families, after its model was given about 4,000 problems. It estimates that the average result took about three hours of compute to produce. OpenAI said some results were formally checked using Lean, a proof assistant, and cautioned that some unformalized results “could have issues.” The source does not provide a full breakdown of which manuscripts received which forms of review.

In software, the report cites figures from several sources. Faros AI reported that teams in periods of high AI adoption merged 98% more pull requests, while review time rose 91%. LinearB, in an analysis of 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. The source also cites a 2026 peer-reviewed study that found 61% of AI-agent pull requests received no human review before they were merged or closed.

A separate example comes from professional services. The report says OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on contracting workflows. In an evaluation across 11 tasks, Astra met 55% of the criteria on average, according to the source. That result indicates improvement over a prior model, but it also leaves a substantial share of evaluation criteria unmet. The material does not give the evaluation methodology or identify the previous model’s score.

At a glance
analysisWhen: The cited mathematical work and softwar…
The developmentA report argues that rising AI-generated output is outpacing the human capacity needed to check and approve it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Is Becoming the Capacity Limit

If organizations can create more work than they can responsibly inspect, the constraint shifts from production to approval. Review time can determine how much AI output is usable, whether a software change is safe to merge, or whether a contract draft is ready for a person to sign. Faster generation alone does not show that the work is accurate or fit for its intended purpose.

The source describes several possible responses already visible in the cited data: some changes may be merged without review, reviewers may deprioritize AI-generated work, or the organization producing the work may decide which items merit outside scrutiny. Each approach has trade-offs. Skipping review can leave errors undetected; treating all AI output as suspect can slow useful work; and relying on the producer’s own selection leaves less independent scrutiny.

The report also raises a workforce concern: review expertise is built through practice. If junior staff spend less time drafting, coding or proving work themselves, they may get fewer opportunities to develop the judgement expected of senior reviewers. Whether AI adoption will produce that effect at scale is not established by the examples provided, but it is a relevant question for employers training the next generation of specialists.

Amazon

AI review tools for software development

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields Show a Similar Gap

The examples span mathematics, software and contract work, but the underlying distinction is similar: automated systems can produce a candidate result, while checking whether it answers the right question may still require people. In mathematics, formal tools can verify that a proof follows from a stated theorem. That does not, by itself, establish that the theorem is the intended one or that the result is important. In software, tests can check the behavior they cover, but cannot guarantee that the tests represent every real requirement.

The source says five leading mathematicians carefully verified an earlier result from the same program: a counterexample to an old Erdős conjecture. It presents that response as a contrast with the volume of newly generated manuscripts. The specific review process and the exact scope of the five mathematicians’ work are not detailed in the supplied material.

Some figures come from companies that sell software review products, including Faros AI and LinearB. That commercial interest is a reason to examine their methods and definitions, not by itself a reason to dismiss their findings. The report says the studies point in a consistent direction, but it does not provide links, full methods or enough detail to independently compare how each source defined review time, acceptance or AI-generated work.

How Much Review Is Being Skipped?

The cited evidence does not establish a single, comparable measure of AI verification across all three fields. The source does not identify the mathematical manuscripts that were formally checked, give full details of the contract evaluation, or provide study methods for every software statistic. The numbers should be understood within their stated samples and definitions, not as universal rates for all organizations or AI tools.

It is also unclear whether increased review times are caused by AI-generated work itself, higher overall submission volume, changes in team practices, or a combination of factors. The report says some reviewers deliberately deprioritize AI-generated changes, but the cited material does not establish how common that practice is beyond the surveyed data. Nor does it show whether organizations are adding review staff or changing processes to address the backlog.

Finally, the proposed “referee premium”—greater value for people able to approve work and take responsibility for it—is an interpretation, not a measured labor-market outcome in the material provided. Evidence about hiring, compensation and training would be needed to determine whether that pattern is already occurring.

Measure Quality Alongside Output

The next test for organizations adopting AI is whether they can track the quality and review status of output, not only how much is produced or how quickly it moves through a workflow. That means making clear what was checked, by whom, against which requirements, and what remains uncertain before work is accepted or used.

For the math results, readers will need clearer information about which manuscripts were formally verified and what independent review they received. For software, further studies with transparent definitions and methods could help distinguish the effect of AI-generated changes from shifts in submission volume or team organization. In contract workflows, published evaluation details would help clarify what the 55% score measures and how performance compares with human review or earlier systems.

The source material does not identify a specific next milestone or timetable. For now, the central issue is whether human review capacity, training and accountability can keep pace with the growing volume of AI-assisted work.

Key Questions

What is the main development described?

The report says AI systems are producing large volumes of mathematical, software and contract-related work, while people remain responsible for checking whether it is correct and suitable for use.

Did OpenAI verify all 722 mathematical manuscripts?

The source does not say that all 722 were formally verified. It says some results were checked in Lean and quotes OpenAI warning that some unformalized results “could have issues.”

What did the cited software data find?

The report cites analyses that found more pull requests in high-AI-adoption periods, longer waits for review of AI-generated changes and some changes receiving no human review. The figures refer to particular datasets and definitions, and some sources sell review tools.

Does the report prove that AI is reducing the supply of expert reviewers?

No. It raises the possibility that less hands-on work for junior staff could weaken the experience through which reviewers develop judgement. The supplied evidence does not establish a measured, economy-wide decline in reviewer supply.

What remains unknown about the contract evaluation?

The source reports that GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks, but does not provide the evaluation method, the previous model’s score or details about how the criteria were weighted.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Spatial Focus Room: Make Distraction Impossible

A new deep-work app for Apple Vision Pro, Spatial Focus Room, removes distractions by physically immersing users in focused environments, redefining concentration.

Circle Launches Arc Mainnet

Circle has launched the Arc mainnet, marking a significant step in its blockchain development. Details are confirmed, but the full scope remains unclear.

The Future Of STEM And AI: ByteDance’s Strategic Investment In AI4S

ByteDance’s new program seeks about 100 researchers for a six-month AI for Science pilot in Beijing, aiming to advance scientific research through AI collaboration.

What AI Achieved In 2026: 9 Major Highlights

A comprehensive overview of the nine most significant advancements in artificial intelligence in 2026, including confirmed breakthroughs and ongoing developments.