LLM生成的解释文本质量高但无实际帮助,反而让人误信。
Quality Without Usefulness: LLM-Generated XAI Narratives as Trust Heuristics Rather Than Decision Aids

- 用LLM把XAI结果转成自然语言,保持高质量但不提升任务表现
- 五项实验显示解释文本不提高准确率,反而虚增使用者信心
- 文本存在感导致错觉,让模型漏掉不可靠预测,适合警惕幻觉者看
先前研究显示,大语言模型(LLMs)可将可解释AI(XAI)输出转化为在可信度、连贯性和可理解性等质量指标上得分较高的自然语言解释(NLEs)。但解释质量是否意味着实际有用?我们在时间序列能源预测领域通过五项受控实验(共60个测试实例,2,730次判断)探究此问题,每项实验对应XAI文献中一种不同的有用性维度。在保持NLE质量处于先前因子实验所确立的高水平前提下,发现NLE在所有五项任务中均未提升任务准确率,反而增加自我报告的信心。安慰剂对照组表明,这种信心提升源于文本存在而非内容本身。在分布外检测任务中,NLE降低了LLM判断者识别不可靠预测的能力,带来虚假安心,掩盖模型失效。我们将其称为‘质量-有用性鸿沟’,主张对XAI到NLE的转化流程评估应超越文本质量指标,延伸至下游任务表现。
原文摘要 · Abstract (English)
Prior work shows that Large Language Models (LLMs) can transform Explainable AI (XAI) outputs into Natural Language Explanations (NLEs) that score highly on quality metrics such as plausibility, coherence, and comprehensibility. But does explanation quality translate to practical usefulness? We investigate this question in a time-series energy forecasting domain through five controlled experiments (2,730 judgments across 60 test instances), each operationalising a distinct facet of usefulness studied in the XAI literature. Holding NLE quality constant at the high levels established by a prior factorial study, we find that NLEs do not improve task accuracy on any of the five tasks, while inflating self-reported confidence. A placebic control shows that this confidence boost is driven by text presence rather than content. In an out-of-distribution detection task, NLEs reduce the LLM judge's ability to flag unreliable predictions, providing false reassurance that masks model failure. We characterise these findings as the Quality-Usefulness Gap and argue that evaluation of the XAI-to-NLE pipeline must extend beyond text-quality metrics to downstream task performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。