arXiv:2604.14325cs.CLcs.AI2026-04被引 1

让大模型的解释更真实,用注意力干预提升决策依据的可信度。

Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance

论文配图:Faithfulness Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance
图 1 · 摘自论文原文
  • 通过注意力干预,引导生成解释时聚焦真实依赖的文本片段。
  • 在多个模型和数据集上,解释的内在忠实性显著提升。
  • 无需训练,适合需要可信解释的高风险应用领域。

大语言模型在自然语言处理中表现卓越,但缺乏可解释性使其被视为黑箱,限制了其在需透明与信任场景的应用。当前主流方法生成看似合理的自然语言解释,但这些解释是否真实反映模型决策所依据的内部证据仍不明确。本文通过反事实分析评估现有解释的内在意图忠实性,发现其往往不忠实。为此提出一种无需训练的方法:利用可信归因方法提取的词级热力图,指导解释生成过程中的注意力分布。该方法在多个模型、基准和提示下均显著提升了解释的内在忠实性。

原文摘要 · Abstract (English)

Large language models (LLMs) achieve strong performance and have revolutionized NLP, but their lack of explainability keeps them treated as black boxes, limiting their use in domains that demand transparency and trust. A promising direction to address this issue is post-hoc text-based explanations, which aim to explain model decisions in natural language. Prior work has focused on generating convincing rationales that appear to be subjectively faithful, but it remains unclear whether these explanations are epistemically faithful, whether they reflect the internal evidence the model actually relied on for its decision. In this paper, we first assess the epistemic faithfulness of LLM-generated explanations via counterfactuals and show that they are often unfaithful. We then introduce a training-free method that enhances faithfulness by guiding explanation generation through attention-level interventions, informed by token-level heatmaps extracted via a faithful attribution method. This method significantly improves epistemic faithfulness across multiple models, benchmarks, and prompts.

可解释性大模型注意力干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。