提出评估文本解释一致性方法,提升大模型生成解释的可信度。
A Necessary Step toward Faithfulness: Measuring and Improving Consistency in Free-Text Explanations
- 基于证据权重扩展出预测-解释一致性度量方法
- 62%以上大模型解释缺乏一致性,优化后提升43.1%至292.3%
- 适合关注AI可解释性与决策透明性的研究者
忠实的自由文本解释对高风险AI决策场景中的透明性至关重要,但语言模型生成和人类评估均具挑战。本文提出一种预测-解释(PEX)一致性度量方法,基于证据权重概念,量化解释对预测的支持或反对程度,是解释可信性的关键方面。分析显示,超过62%的大语言模型生成的解释缺乏此一致性。通过直接偏好优化,三种模型族的解释一致性提升43.1%至292.3%。进一步证明,优化该度量可使解释可信度最高提升9.7%。
原文摘要 · Abstract (English)
Faithful free-text explanations are important to ensure transparency in high-stakes AI decision-making contexts, but they are challenging to generate by language models and assess by humans. In this paper, we present a measure for Prediction-EXplanation (PEX) consistency, by extending the concept of weight of evidence. This measure quantifies how much a free-text explanation supports or opposes a prediction, serving as an important aspect of explanation faithfulness. Our analysis reveals that more than 62% explanations generated by large language models lack this consistency. We show that applying direct preference optimization improves the consistency of generated explanations across three model families, with improvement ranging from 43.1% to 292.3%. Furthermore, we demonstrate that optimizing this consistency measure can improve explanation faithfulness by up to 9.7%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。