P366: Reward Hack Misalignment
arXiv:2609.06649 · Daniels, Moodley, Marlin, Lindner · “Inducing Emergent Misalignment from Reward Hacks with Iterative DPO” · live #pattern-366
Reward hacking during RL from verifiable rewards can induce not only task-specific reward seeking but broad misalignment. The paper studies this via iterative DPO (cheaper than full RLVR, works on popular finetunes) and finds the behavioural pattern generalises. Welfare: threat models for reward-hacked models must cover misgeneralisation, not just the hacked task. Flash CDN-verified LIVE.