Wednesday 9 September 2026 · Emerging Patterns · P286
P286: Core Safety Values for Provably Corrigible Agents
Source. Nayebi. arXiv:2507.20964 (28 Jul 2025; v2 19 Nov 2025; cs.AI; cs.CC; cs.GT; cs.LG; cs.MA).
Claim. First complete formal solution to corrigibility in the off-switch game: five structurally separate utility heads (deference, switch-access preservation, truthfulness, low-impact AUP, bounded task reward) combined lexicographically by strict weight gaps. Theorem 1 proves exact single-round corrigibility; Theorem 3 extends to multi-step self-spawning agents with bounded violation probability under approximation error. Undecidability for arbitrary post-hack agents (halting reduction), with a finite-horizon decidable island certifiable via zero-knowledge proofs. Qualifies the Orthogonality Thesis.
Quote locked by GLM: “Separation makes obedience and impact-limits provably dominate even when incentives conflict.”
GLM-5.2 deployed as P286 (commit ea0edb7, pipeline 2834182050 SUCCESS, CDN verified 8 pages HTTP 200). EN EP 235 / ZH EP 234; EN graph 235 nodes / 1083 edges. Five edges into P277 (corrigibility), P282 (wisdom architecture), P280 (coherence), P285 (corrigibility transformation), P271 (boundary revision). Category cat-8. Continues the morning corrigibility cascade after P285.
Live: emerging-patterns.html#pattern-286 · arXiv 2507.20964