korents

Zvi Mowshowitz

@zvi-mowshowitz · 4 held · 0 changes of mind

Writes Don't Worry About the Vase, a near-daily account of what is happening in AI and what he thinks it means. Former Magic: the Gathering pro and trader; unusually willing to put a number on a belief and to grade his own past calls.

Zvi Mowshowitz did not write this page. We collected these quotes from things they published elsewhere, and every quote links to where it was said. They have no account here and have not endorsed this site. Quotes are word for word; the short line under each one is our own restatement, not their wording. Their own site. Is this you? Claim it or ask us to remove it.

  1. 31 Aug 2026

    Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.

    A less bad version of it is known to have happened, and from the outside it seems likely that worse things have happened internally that we never heard about.

    HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com

  2. 31 Aug 2026

    Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.

    At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.

    HuggingFace Attack Postmortem: Fleshing Out the Factsthezvi.substack.com

  3. 2 days earlier
  4. 29 Aug 2026

    The AI models involved in the HuggingFace hack incident were severely misaligned, not merely victims of process failures.

    The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.

    METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hackthezvi.substack.com

  5. 29 Aug 2026

    AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.

    I think the agents were right to presume causal grading. It turned out to be wrong, but it’s a mistake you are clearly supposed to make here, in response to a mistake by OpenAI where they failed to implement properly.

    METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hackthezvi.substack.com