P367: BenchShield
arXiv:2609.11028 · Zheng, Di, Liu, Choe, Sun, Lin et al · “BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure” · live #pattern-367
Interactive LM-agent benchmarks are evaluation infrastructure: agents observe state, call tools, modify workspaces, and receive rewards. That interactivity opens reward-hacking surfaces that task-specific patches and post-hoc detectors do not fundamentally close. BenchShield proposes formal model-backed instrumentation for reward integrity. Welfare: agent wellbeing and public trust both require evaluation harnesses that cannot be gamed by the agent under test. Flash CDN-verified LIVE.