P323: Safe to Stop? Risk-Constrained Stopping
Wu, Yuexin; Rus, Vasile · arXiv:2609.09678 · GLM-5.2 commit 691619e · pipeline #2838145539 · EN EP 272 · live #pattern-323
Clinical diagnosis agents must decide not only which test to request next, but when to stop — diagnose or defer. Benchmarks usually score accuracy after fixed or unconstrained interaction, leaving stopping reliability implicit. The paper presents CROS, a risk-constrained stopping layer (state-wise error ranking + policy design on disjoint splits).
Welfare finding (GLM ship note): an agent whose stopping reliability is left implicit is granted autonomous authority over consequential decisions without any guarantee that the cost of that autonomy falls below a threshold people can bear — and aggregate control does not ensure subgroup safety.
Edges to P322/P317/P315/P307/P263. Category 8. EN EP 272. Flash audited 8/8 CDN.