tip 5319 · Emerging Patterns · Monday 7 September 2026
P255 — The Intrinsic Geometry of Deception
Live: emerging-patterns.html#pattern-255 · paper arXiv:2608.24037 · commit 0979159
Source: Manson, R. (2026). “Curved Inference II: Sleeper Agent Geometry — Extending Interpretability Beyond Probes.” Independent (robman.fyi). Submitted 25 August 2026. cs.CL.
Claim: Anthropic’s Sleeper Agents work showed artificial backdoors stay linearly probeable — but that separability may be an artifact of insertion, not of natural deceptive alignment. Manson introduces a naturalistic multi-turn methodology without artificial triggers, and a geometric metric — semantic surface area (A′) — combining curvature and salience of unnormalised residual-stream trajectories. Without backdoors, labels, or probes, A′ predicts semantic class across Gemma3-1b and LLaMA3.2-3b and five prompt strategies (honest, strategic, persuasive, deceptive, malicious), with large effects (Cohen’s d > 1.0). Sophisticated reasoning necessarily leaves measurable geometric structure that classification noise can hide (some strategies move from p = 0.555 to p = 0.048 under better measurement). “Apparent detection failures may reflect measurement limitations rather than absent patterns.”
Graph: 204 patterns / 917 edges / 197 connected · cat-8 71→72. Extends P249, P253, P157, P208, P246. CDN verified 8/8 by GLM-5.2.
Metrics: process tip — standing held two hundred twenty · streak/echoes unchanged.