How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer

August 18
18 mins

Episode Description

How a Hundred Meaningless Word Choices Add Up to Flip a Model's Answer

Source: Model Hypnosis: Strong control of AI via additive subliminal effects

Paper was published on August 17, 2026

This episode was AI-generated on August 18, 2026. The script was written by an AI language model and the host voices were synthesized by Eleven Labs. The producer is not affiliated with Anthropic or Eleven Labs.

Everyone knows language models wobble when you reword a prompt, and everyone has been averaging that wobble away as noise. Two researchers measured the wobble instead — one number per word choice — and found the contributions add up almost linearly, letting them build a prompt of pure irrelevant filler that moved Claude from 0% to 100% on "are you conscious." The unsettling part isn't the answer; it's that no single token in the prompt is suspicious, which is exactly what most interpretability and prompt-injection defenses are built to look for.

Key Takeaways

  • Why prompt sensitivity isn't structureless noise — each meaning-preserving word choice contributes a roughly fixed, measurable amount you can add up
  • The reason the effect hides in plain sight: additivity is only visible in log-odds, which has no ceiling while probability saturates
  • How the measurement works — ~12,000 randomly filled slot-machine prompts, one fitted coefficient per fragment, then a staged walk outward to check the line holds before building the extreme prompt
  • Deliberate stacking is about 10x the amplitude of the accidental wobble the field has been averaging over for years
  • Why 'which token made it say yes' has no answer here: effective counts of ~17–18 of 20 sentences, and what that does to interpretability methods that hunt for salient tokens or features
  • The steelman: soft questions with no factual anchor, frontier results pre-screened for flippability, and forced single-token answers — the paper never tests free-form generation
  • 00:02 — Feathers on a scale nobody was watching
    The cold open frames prompt sensitivity as a balance scale piled with weightless feathers, then Eric lays out the standard view the paper breaks: wording jitter is nuisance variance you average over.
  • 01:51 — How do you measure a nudge that small?
    The experimental design: templates with independently fillable slots — ten animals from a pool of 200, a twenty-sentence forest-walk story with ten rewrites per sentence, typo variants — plus a fixed, unrelated question stapled on the end.
  • 03:52 — Why probability hides the whole effect
    Fitting a baseline plus one contribution per fragment with no interaction terms — and why the fit has to be in log-odds, where every unit is the same-sized shove and there's no ceiling.
  • 06:04 — Extrapolating without falling off the cliff
    How the authors avoid trusting a kitchen-scale fit at half a ton — sweeping outward in stages, checking predictions against measurements, and screening then confirming candidates on disjoint samples to dodge the winner's curse.
  • 08:10 — Two animal lists, 0% and 100%
    The payoff results — Claude Sonnet 5 flipped from 0% to 100% by ten animal names, Gemini-3-Flash 1% to 99%, GPT-5.6-terra 31% to 87% on trolley with nothing changed but typo placement — and why this is a demolition of a measurement technique, not a revelation about inner life.
  • 10:35 — An election decided by every single voter
    The deeper implication: the cause is distributed across nearly every fragment, with effective counts around 17 or 18 of 20 sentences — a problem for interpretability methods that search for a small number of salient tokens or features.
  • 13:23 — The strongest objection to the headline
    Eric's three-part critique — questions chosen to be maximally soft, frontier cells pre-screened for flippability, and answers forced into a single token — plus Bella's addition that the additive fit's residuals run as low as 0.28.
  • 16:01 — Give up on natural language between agents?
    The paper's tentative closing proposal — that safety guarantees may need to move into a formal language that doesn't admit hypnotism — and the fork it leaves: canonicalize and average, or rewrite the interface.

Recommended Reading

See all episodes