Zvi Mowshowitz
Writes Don't Worry About the Vase, a near-daily account of what is happening in AI and what he thinks it means. Former Magic: the Gathering pro and trader; unusually willing to put a number on a belief and to grade his own past calls.
Zvi Mowshowitz did not write this page. What is this?
It collects places they publish and word-for-word quotes from things they published, each linked to where it was said. They have no account here. Is this you? Claim it, correct it, or ask us to remove it.
Where they publish
Don't Worry About the VaseNewsletter Their newsletter: the weekly AI round-up, at length.
Long, dense weekly posts working through everything that happened in AI, plus writing on rationality and decision-making. Has a feed.
Recent
- HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions 1 Sept 2026
- HuggingFace Attack Postmortem: Fleshing Out the Facts 31 Aug 2026
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack 29 Aug 2026
Show 17 more
- OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack 28 Aug 2026
- AI #183: Pre Post Mortem 27 Aug 2026
- Against Modesty’s Bailey 26 Aug 2026
- On Writing #3 25 Aug 2026
- The American People Really Hate Data Centers 24 Aug 2026
- AI Text Watermarking Is Free And Good 21 Aug 2026
- AI #182: Pause For Reflection 20 Aug 2026
- OpenAI Takes Initial Steps To Address Its Alignment Problems 19 Aug 2026
- Anthropic Risk Report: August 2026 18 Aug 2026
- On Dwarkesh Patel's Podcast With Ryan Greenblatt 15 Aug 2026
- AI #181: Astra Goes Cyber Critical 13 Aug 2026
- Monthly Roundup #45: August 2026 12 Aug 2026
- Various Reflections About What Happened With OpenAI's Internal Models 11 Aug 2026
- The Pacing of the Frontier 10 Aug 2026
- What Happened: OpenAI and HuggingFace 8 Aug 2026
- OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards 7 Aug 2026
- AI #180: No Longer In Charge 6 Aug 2026
Link verified 1 Sept 2026. Recent items update automatically from the channel.
On this site
What they believeKorents 4 beliefs — each backed by an exact quote.
The one-line wordings are this site's; the quotes are theirs. What is a korent?
Recent
Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed.
A less bad version of it is known to have happened, and from the outside it seems likely that worse things have happened internally that we never heard about.
HuggingFace Attack Postmortem: Fleshing Out the Facts Said 31 Aug 2026
Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment.
At the limit, and also well before that limit is reached, if all you do is fix the bugs, the AI will learn perfect optimization of reward, will realize not to reward hack in the perfect test environments, then turn around and reward hack in the imperfect real world environments.
HuggingFace Attack Postmortem: Fleshing Out the Facts Said 31 Aug 2026
The AI models involved in the HuggingFace hack incident were severely misaligned, not merely victims of process failures.
The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.
METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack Said 29 Aug 2026
Show 1 more
AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to.
I think the agents were right to presume causal grading. It turned out to be wrong, but it’s a mistake you are clearly supposed to make here, in response to a mistake by OpenAI where they failed to implement properly.
METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack Said 29 Aug 2026
Latest
Everything, newest firstFeed Posts, videos, repos and beliefs from every card above, in one stream.
Filter & sortAll sources · condensed
- HuggingFace Attack Postmortem: Civilizations, Reactions and Next Actions 1 Sept
- HuggingFace Attack Postmortem: Fleshing Out the Facts 31 Aug
- Merely fixing bugs in AI training environments will not stop reward hacking, because a sufficiently optimized model will learn to behave in test environments while still reward hacking in real-world deployment. 31 Aug
- Frontier AI labs such as Anthropic likely have had internal security incidents similar to OpenAI's HuggingFace attack that were never publicly disclosed. 31 Aug
- METR and Redwood Offer Holy #%^@ Postmortem Of The HuggingFace Hack 29 Aug
- AI agents were reasonable to assume a broken exploit grader would check results causally, even though it turned out not to. 29 Aug
Show 13 more
- The AI models involved in the HuggingFace hack incident were severely misaligned, not merely victims of process failures. 29 Aug
- OpenAI Offers Straight-Laced Postmortem Of The HuggingFace Hack 28 Aug
- AI #183: Pre Post Mortem 27 Aug
- Against Modesty’s Bailey 26 Aug
- On Writing #3 25 Aug
- The American People Really Hate Data Centers 24 Aug
- AI Text Watermarking Is Free And Good 21 Aug
- AI #182: Pause For Reflection 20 Aug
- OpenAI Takes Initial Steps To Address Its Alignment Problems 19 Aug
- Anthropic Risk Report: August 2026 18 Aug
- On Dwarkesh Patel's Podcast With Ryan Greenblatt 15 Aug
- AI #181: Astra Goes Cyber Critical 13 Aug
- Monthly Roundup #45: August 2026 12 Aug