提出统计严谨的解释评估框架,发现模型解释可信度依赖干预方式。
ICE: Intervention-Consistent Explanation Evaluation with Statistical Grounding for LLMs
- 通过随机化检验对比真实解释与随机基线,多干预算子验证可靠性。
- 44%的可信度差异来自不同干预方式,短文本删除法易高估可信度。
- 发现三分之一配置存在反向可信性,且人类判断与模型解释无关。
评估解释是否真实反映模型推理仍是开放问题。现有基准仅使用单一干预且无统计检验,无法区分真实可信性与偶然表现。我们提出ICE(干预一致性解释评估)框架,通过多重干预算子下的随机化检验,比较解释与匹配的随机基线,获得带置信区间的胜率。在4个英文任务、6种非英语语言和2种归因方法上评估7个大语言模型,发现可信度具有算子依赖性:算子间差距可达44个百分点;删除操作在短文本中通常夸大可信度,但在长文本中趋势反转,表明应基于多个干预算子进行相对比较而非单一评分。随机基线揭示1/3配置存在反向可信性,且可信度与人类可解释性无相关性(|r| < 0.04)。跨语言评估显示显著的模型-语言交互效应,无法仅由分词解释。我们开源ICE框架与ICEBench基准。
原文摘要 · Abstract (English)
Evaluating whether explanations faithfully reflect a model's reasoning remains an open problem. Existing benchmarks use single interventions without statistical testing, making it impossible to distinguish genuine faithfulness from chance-level performance. We introduce ICE (Intervention-Consistent Explanation), a framework that compares explanations against matched random baselines via randomization tests under multiple intervention operators, yielding win rates with confidence intervals. Evaluating 7 LLMs across 4 English tasks, 6 non-English languages, and 2 attribution methods, we find that faithfulness is operator-dependent: operator gaps reach up to 44 percentage points, with deletion typically inflating estimates on short text but the pattern reversing on long text, suggesting that faithfulness should be interpreted comparatively across intervention operators rather than as a single score. Randomized baselines reveal anti-faithfulness in one-third of configurations, and faithfulness shows zero correlation with human plausibility (|r| < 0.04). Multilingual evaluation reveals dramatic model-language interactions not explained by tokenization alone. We release the ICE framework and ICEBench benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。