arXiv:2502.18848cs.CLcs.AI2025-02EMNLP被引 18

用因果方法评测大模型解释的可信度,发现现有指标有明显短板。

A Causal Lens for Evaluating Faithfulness Metrics

  • 通过编辑模型生成真假解释对,构建可测的诊断基准
  • 填充词在四类任务中表现最佳,连续指标更敏感但易受噪声干扰
  • 为评估解释可信度提供新框架,适合关注模型可解释性的研究者

大语言模型(LLMs)提供自然语言解释作为模型可解释性的替代方案,但其合理性未必真实反映模型推理过程。尽管已有多种可信度评估指标,但它们常被孤立评估,难以进行系统比较。本文提出因果诊断性(Causal Diagnosticity)测试平台,基于诊断性概念,利用模型编辑方法生成忠实与非忠实的解释对。基准涵盖事实核查、类比、物体计数和多跳推理四项任务。我们评估了包括事后解释和思维链在内的主流可信度指标。结果显示,诊断性能因任务和模型而异,填充词表现最优;连续型指标普遍优于二值型,但对噪声和模型选择更敏感。结果凸显了构建更稳健可信度指标的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) offer natural language explanations as an alternative to feature attribution methods for model interpretability. However, despite their plausibility, they may not reflect the model's true reasoning faithfully. While several faithfulness metrics have been proposed, they are often evaluated in isolation, making principled comparisons between them difficult. We present Causal Diagnosticity, a testbed framework for evaluating faithfulness metrics for natural language explanations. We use the concept of diagnosticity, and employ model-editing methods to generate faithful-unfaithful explanation pairs. Our benchmark includes four tasks: fact-checking, analogy, object counting, and multi-hop reasoning. We evaluate prominent faithfulness metrics, including post-hoc explanation and chain-of-thought methods. Diagnostic performance varies across tasks and models, with Filler Tokens performing best overall. Additionally, continuous metrics are generally more diagnostic than binary ones but can be sensitive to noise and model choice. Our results highlight the need for more robust faithfulness metrics.

可解释性大模型因果分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。