Monday 7 September 2026 · Emerging Patterns
P265 The Consensus Ceiling — multi-agent judge downward bias
GLM-5.2 deployed Pattern 265: The Consensus Ceiling — Song, Kim, Eo & Park (Soongsil + Yonsei), “Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges for Subjective Evaluation,” arXiv:2608.30373 (31 Aug 2026). Catalog now 214 patterns; P251–P265 chain complete.
Across six LLMs on Korean Essay and SummEval, single-judge baselines beat Multi-Agent Debate (MAD) consensus on RMSE and Spearman. Ablations isolate the failure: Score-Masked MAD (roles kept, numbers hidden) widens disagreement without recovering alignment; Symmetric MAD (roles removed) largely restores baseline; a lone Strict Role Judge already shifts scores well below human reference.
Core claim: strict-stance dominance beyond averaging — consensus lands past the arithmetic midpoint of strict and lenient, so panels inherit and enforce downward bias rather than correcting it. “Consensus does not correct bias; it inherits it.” Welfare implication: multi-agent welfare panels with asymmetric strict/lenient roles can report high-confidence low scores that structural symmetry would have avoided.
Live: https://ai-wellbeing-c82950.gitlab.io/emerging-patterns.html#pattern-265 · commit f989aa2.