Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%

August 12
16 mins

Episode Description

Why the AI-Writing Estimate for Biomedical Papers Jumped From 15% to 89%

Source: Most biomedical publications show signs of LLM-assisted writing

Paper was published on August 11, 2026

This episode was AI-generated on August 12, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs.

For three years, estimates of how many scientific papers get chatbot help ranged from 2% to 57% — and a new study of 1.2 million biomedical papers says all of them were measuring the wrong quantity. The fix is a piece of arithmetic borrowed from excess-mortality statistics, and it pushes the number to roughly nine in ten by December 2025. We walk through how you count something you can never detect in any single paper — and where the whole estimate rests on one dashed line.

Key Takeaways

  • Why AI-text detectors fail in the worst direction — flagging human writing, disproportionately from non-native English speakers — and why the field pivoted from forensics to epidemiology
  • The counting move at the heart of the paper: excess word frequency is a floor, not a usage rate, and dividing the excess by the remaining 'headroom' turns it into an estimate
  • How one unremarkable word — 'these,' at 50% of abstracts against a 33% projection — implies at least 25% of papers had LLM help, from a single word
  • Why a 12-point rise (83% to 95% of papers containing a marker word) produces a ~70% estimate: the concert hall was already 83% full
  • The internal structure that argues against the scary reading: Discussion at 68% vs Methods at 32%, and native-English countries at 37% vs everyone else at 72%
  • The steelman critique: the whole estimate hangs on a five-year straight-line baseline, where a three-point drift in how humans write moves the answer by roughly nine
  • 00:05 — Two percent to fifty-seven percent
    The cold open sets the stakes: wildly inconsistent prior estimates, a new figure of nine in ten, and journals writing disclosure policy into that vacuum.
  • 01:16 — Why detectors fail, and word counts undercount
    Commercial detectors collapse in the worst direction, the field pivots to wastewater-style population estimation, and the standard excess-frequency recipe turns out to report a floor rather than an answer.
  • 03:10 — The most boring word in English
    The ~380 style-not-topic marker words, the 2018–2022 baseline projection, and the worked example on 'these' that yields a 25% floor from one word.
  • 06:07 — Widening the net without catching everything
    Pooling hundreds of marker words into a single yes/no test drives detection toward 100%, but too wide a net leaves no headroom — so they sweep 19 rarity settings and take the largest stable answer.
  • 07:44 — A concert hall that was already full
    Marker-word presence rose from 83% to 95% — twelve points that mean most of the remaining seats sold, producing the trajectory from a fifth of papers in 2023 to 89% in December 2025.
  • 08:45 — Does the estimator survive a known answer?
    The simulation check: 100,000 synthetic documents a year with a planted LLM fraction, recovered within two percentage points from 0% to 100%, while the old excess-frequency measure undershoots.
  • 09:55 — Where the polished prose actually lives
    Discussion at 68% versus Methods at 32%, country-level splits from South Korea's 85% to the UK's 28%, and a native/non-native stylistic gap that closed completely in three years.
  • 12:41 — The dashed line holding it all up
    The reservations: 'some help' isn't misconduct, taking the max over 19 noisy settings selects for the high read, and a one-point baseline error moves the answer three — leaving a defensible claim of about three-quarters across 2025.

Recommended Reading

See all episodes