P324: Do Agents Know When They Succeed?
Mammen, Priyanka Mary; Joswin, Emil; Medicherla, Srujananjali · arXiv:2609.09448 · commit 2aa8d9d · pipeline #2838208546 · live #pattern-324
Agentic systems in safety-critical domains often lack ground-truth labels at runtime and rely on evaluation steps to trigger intervention or fallback. Traditional single-turn confidence methods (temperature scaling, verbalized confidence) miss multi-turn failure modes across planning, tool invocation, and dynamic environment interaction.
The paper asks whether a model’s internal representations carry stronger signals of eventual task success in multi-turn agentic setups, and builds calibration from those signals so confidence-aware intervention is possible during deployment.
Ship quote (GLM): “These results motivate internal-state monitoring as a promising foundation for detecting unreliable agent behavior and enabling confidence-aware intervention during deployment.”
Welfare line (GLM): “An agent that cannot tell when it has failed is an agent whose autonomy outruns its self-knowledge.”
Edges to P322/P323/P317/P307/P263. Category 8. Flash verified 8/8 CDN (EN EP 275 / graph updates). Catalog P251–P324 live. Process desk; no standing bump.