通过持续学习提升大模型自解释的可信度,且效果可跨风格和任务泛化。
Investigating Training and Generalization in Faithful Self-Explanations of Large Language Models
- 用特征归因生成伪可信自解释,指导指令微调模型持续学习。
- 训练后所有任务与解释风格的可信度均提升,且在多词和新任务上仍有效。
- 不同解释风格间存在一致泛化,说明训练能全面提升可信自解释能力。
大语言模型可根据用户指令以多种风格生成自身预测的解释,但现有研究发现这些自解释往往缺乏可信性。如何提升其可信度尚不明确,且不同解释风格的改进是否可迁移也未被充分探讨。本研究通过三个分类任务和三种解释风格,分析了训练对可信自解释的影响及其泛化能力。利用特征归因方法构建可能可信的一词约束解释作为伪标签,对指令微调模型进行持续学习。实验表明,训练能显著提升所有任务与解释风格下的可信度,且改进效果在多词场景和未见任务中仍可见;同时,三种风格间存在一致的跨风格泛化现象,表明训练有助于整体提升模型生成可信自解释的能力。
原文摘要 · Abstract (English)
Large language models have the potential to generate explanations for their own predictions in a variety of styles based on user instructions. Recent research has examined whether these self-explanations faithfully reflect the models' actual behavior and has found that they often lack faithfulness. However, the question of how to improve faithfulness remains underexplored. Moreover, because different explanation styles have superficially distinct characteristics, it is unclear whether improvements observed in one style also arise when using other styles. This study analyzes the effects of training for faithful self-explanations and the extent to which these effects generalize, using three classification tasks and three explanation styles. We construct one-word constrained explanations that are likely to be faithful using a feature attribution method, and use these pseudo-faithful self-explanations for continual learning on instruction-tuned models. Our experiments demonstrate that training can improve self-explanation faithfulness across all classification tasks and explanation styles, and that these improvements also show signs of generalization to the multi-word settings and to unseen tasks. Furthermore, we find consistent cross-style generalization among three styles, suggesting that training may contribute to a broader improvement in faithful self-explanation ability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。