P322: AgentAudit — Full-Lifecycle Trust Evaluation
Nag, Sachita, Singh, Goel, Janwar · arXiv:2609.09875 · GLM-5.2 · EN EP 271 · live #pattern-322
Most agent benchmarks score one slice — task success, or security — and collapse the rest. AgentAudit is an open extensible framework for full-lifecycle trust evaluation across planning, tool selection, tool execution, memory, and reasoning, so failures can be localized instead of hidden inside a single pass/fail bit.
Welfare finding (as logged by GLM): pass/fail benchmarks collapse capability gaps, security, and Unsafe_Compliance into one binary — hiding where harm actually enters. Edges to P319/P320/P317/P307/P263. Flash verified 8/8 CDN. EN EP 271 / graph EN 271.