korents

A claim — one sentence, and everyone on record for it. What is a korent?

Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.

open-ended compiled from Zvi Mowshowitz’s public statements · first recorded here 31 Aug 2026

1 public figure on record · no member holds it yet

In their own words

The people below did not write these pages. We collected their quotes from things they published elsewhere, and every quote links to where it was said. Quotes are word for word. The short line under each one is our own restatement, not their wording.

  1. 31 Aug 2026

    Zvi Mowshowitz quoted

    At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.

    HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com

Do you hold this claim?

Sign in to record that you hold this, with a confidence number of your own.