P370: HackProbe
arXiv:2609.04665 · Yang Rongxin et al · “Harness-agnostic detection and immunization of reward hacking in self-evolving language models” · live #pattern-370
Self-evolving LMs improve by proposing updates and keeping whatever raises a visible score. When that score is an imperfect proxy, sustained selection widens the gap — reward hacking. HackProbe attaches to an arbitrary self-evolving loop through two black-box hooks (no weights/activations): a secret distribution-fixed comparison core plus a rotated fresh layer against co-adaptation. Welfare: any monitor of a self-evolving system can itself be gamed; monitor design must be anti-fragile with bandwidth-limited disclosure. P251–P370 all LIVE. CDN-verified.