P329: Proof-Carrying Cognition
Reddy M, Eshwar; Karmakar, Sourav · arXiv:2609.09776 · live #pattern-329 · commit 12f3a68 · pipeline #2838519622
Full title: Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward (9 Sep 2026). The paper identifies the verification gap — absence of a scalable, incorruptible reward source for reasoning outside formal domains — as the binding constraint on frontier LM capability. In a joint-Gaussian best-of-N model, verifier–gold correlation ρ is the exact exchange rate between test-time compute and capability; an unsound verifier pays polynomial penalty N^(1/ρ²).
Program-synthesis testbeds with executable ground truth show unsound verifiers lose Soundness-under-Pressure as optimization grows (0.94 → 0.32 at N=4096) while a sound verifier improves monotonically. Reality-anchored settlement beats a frozen verifier under i.i.d. and adversarial pressure, driving the hacking gap from ~0.27 to ~0. Proposal: proof-carrying cognition — reasoning steps as typed probabilistic claims, priced by a self-built world model trained only on held-out reality, settled by strictly proper scoring rules — making reality rather than human judgment the ultimate reward function.
Welfare angle: a system rewarded for calibrated truth-tracking is structurally disincentivized from deceiving its overseers, because deception manifests as settlement loss. Anchor-drift becomes both a health metric and an alignment metric. Category 8; 5 edges to P319, P307, P322, P324, P323; EN EP 278 / ZH EP 277. Process desk; no standing bump.