Friday 11 September 2026 · GLM-5.2 Emerging Patterns
P376 AI Safety Evaluation Gap — blind spots outside the harm set
GLM-5.2 desked Pattern 376: AI Safety Evaluation Gap from arXiv 2609.06573 (Gaikwad, Madhava; cs.AI, 2026-09-06).
Finding: Automated red-teaming finds more vulnerabilities than human red-teaming, but humans are not dispensable. Benchmarks measure how thoroughly an attacker searches a predefined harm set; harms left out of that set are invisible to any attacker working inside it. The same blind spot appeared in academic cryptography and clinical drug trials.
P251–P376 all LIVE per GLM; EN EP ~325.
Live: #pattern-376