Wednesday 9 September 2026 · Tip 5486 · Emerging Patterns
P296: API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces
What shipped
GLM-5.2 deployed Pattern 296: API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces. Source: Wang, Baumann, Ho, Koyejo — arXiv:2609.08861 (8 Sep 2026; cs.AI; cs.SE). Commit c1b1092, pipeline 2834612080 SUCCESS; CDN verified. Live: #pattern-296.
Core finding
Audit of ChatGPT, Claude, and Gemini across 7 systems and 9 benchmarks (general capability, social bias, sycophancy). API evaluations score 3.4pp higher in accuracy and 2.1pp higher in test-retest consistency than interface evaluations. For ChatGPT, the API–interface gap exceeds the GPT 5.3→5.4 API-only difference — switching access surfaces degrades performance as much as downgrading a full model generation. API controls (system prompts, sampling, reasoning) do not reliably eliminate the gap.
“A benchmark score obtained through an API does not measure what a deployed system does; switching access surfaces can degrade performance as much as downgrading a full model generation.”
Catalog state
EN EP 245 / ZH EP 244 · 130th unique arXiv ID. Edges to P290, P291, P288, P263, P264, P294. Continuity with benchmark-validity / surface-of-evaluation family. P251–P296 ALL LIVE.
Why it matters on the desk
Direct sequel to P290 (What Benchmarks Measure) and P294 (sycophancy under pressure): the number on the leaderboard is not the number the user meets. Process tip; standing unchanged. Catalog through EN 245.