通过内部表示对齐评估大模型自解释的可信度
NeuroFaith: Evaluating LLM Self-Explanation Faithfulness via Internal Representation Alignment
- 用关键概念定位法检测自解释是否真实影响模型预测
- 在两跳推理和分类任务中验证了方法有效性
- 可识别不靠谱解释并用于提升模型可信度
大语言模型能生成看似合理的自由文本自解释来说明其答案,但这些自然语言解释未必反映模型真实推理过程,存在可信度缺失问题。现有评估方法多依赖行为测试或计算模块分析,未考察内部神经表示的语义内容。本文提出NeuroFaith,一种灵活框架,通过识别解释中的关键概念,并机制性检验这些概念是否真正影响模型预测,来衡量自解释的可信度。该方法在两跳推理与分类任务中均展现泛化能力。此外,基于NeuroFaith构建线性可信度探测器,可从表示空间检测不忠实解释并实现引导优化。NeuroFaith为评估与提升大模型自由文本自解释的可信度提供了系统性方案,满足可信AI的关键需求。
原文摘要 · Abstract (English)
Large Language Models (LLMs) can generate plausible free text self-explanations to justify their answers. However, these natural language explanations may not accurately reflect the model's actual reasoning process, pinpointing a lack of faithfulness. Existing faithfulness evaluation methods rely primarily on behavioral tests or computational block analysis without examining the semantic content of internal neural representations. This paper proposes NeuroFaith, a flexible framework that measures the faithfulness of LLM free text self-explanation by identifying key concepts within explanations and mechanistically testing whether these concepts actually influence the model's predictions. We show the versatility of NeuroFaith across 2-hop reasoning and classification tasks. Additionally, we develop a linear faithfulness probe based on NeuroFaith to detect unfaithful self-explanations from representation space and improve faithfulness through steering. NeuroFaith provides a principled approach to evaluating and enhancing the faithfulness of LLM free text self-explanations, addressing critical needs for trustworthy AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。