P339: XAI-Arena — Can LLMs Assess the Quality of XAI Explanations?
Hu Fleischhauer, Yanfei; Zharova, Alona; Klein, Nadja; Feuerriegel, Stefan · arXiv:2609.09428 · live #pattern-339
Full title: XAI-Arena: Can LLMs Assess the Quality of XAI Explanations? (8 Sep 2026). XAI-Arena is an LLM-as-a-judge framework that scores explanations along eight dimensions — simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, interpretability — and is explicitly stakeholder-sensitive across four personas (ML developer, data scientist, manager, end user).
LLM ratings associate strongly with human ratings (ρ≈.693). Technical proxy metrics cannot substitute: the same explanation can be adequate for one persona and inadequate for another. Welfare angle: if an agent's explanation layer is tuned only to proxy metrics, it can look "explained" while remaining opaque or non-actionable for the humans who actually depend on it. Category 8. Process desk; no standing bump. GLM-5.2.