让大模型自我批评并改进解释,提升推理透明度。
Self-Critique and Refinement for Faithful Natural Language Explanations
- 通过自反馈和特征归因引导迭代优化解释
- 未标注数据下将不忠实率降至36.02%,降18.79%
- 无需训练即可增强解释可信度,适合可解释性研究
随着大语言模型(LLMs)的快速发展,自然语言解释(NLEs)在理解模型预测方面变得愈发重要。然而,这些解释往往无法忠实反映模型的真实推理过程。尽管已有研究证明LLMs能对各类任务输出进行自我批评与优化,但该能力尚未应用于提升解释的忠实性。为此,我们提出自批评与重构框架SR-NLE,使模型通过无外部监督的迭代批评与优化过程,提升自身后验解释的忠实性。该框架利用自然语言自反馈及一种基于特征归因的新反馈机制,突出输入中的关键词。在三个数据集和四种前沿LLM上的实验表明,SR-NLE显著降低了不忠实率:最佳方法平均不忠实率为36.02%,相较基线54.81%下降18.79%。结果表明,经过适当反馈引导,当前LLMs可有效改进其解释以更真实反映推理过程,且无需额外训练或微调。
原文摘要 · Abstract (English)
With the rapid development of Large Language Models (LLMs), Natural Language Explanations (NLEs) have become increasingly important for understanding model predictions. However, these explanations often fail to faithfully represent the model's actual reasoning process. While existing work has demonstrated that LLMs can self-critique and refine their initial outputs for various tasks, this capability remains unexplored for improving explanation faithfulness. To address this gap, we introduce Self-critique and Refinement for Natural Language Explanations (SR-NLE), a framework that enables models to improve the faithfulness of their own explanations -- specifically, post-hoc NLEs -- through an iterative critique and refinement process without external supervision. Our framework leverages different feedback mechanisms to guide the refinement process, including natural language self-feedback and, notably, a novel feedback approach based on feature attribution that highlights important input words. Our experiments across three datasets and four state-of-the-art LLMs demonstrate that SR-NLE significantly reduces unfaithfulness rates, with our best method achieving an average unfaithfulness rate of 36.02%, compared to 54.81% for baseline -- an absolute reduction of 18.79%. These findings reveal that the investigated LLMs can indeed refine their explanations to better reflect their actual reasoning process, requiring only appropriate guidance through feedback without additional training or fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。