用小模型生成易懂的反事实解释,让AI决策更透明。
Enhancing XAI Narratives through Multi-Narrative Refinement and Knowledge Distillation
- 用大中小语言模型协作生成解释叙事,提升可读性。
- 小模型经知识蒸馏后推理能力接近大模型。
- 新评估方法验证解释是否符合真实反事实逻辑。
可解释人工智能(XAI)致力于揭示深度学习模型的决策过程。其中,反事实解释因其能通过最小改动揭示预测变化而备受关注。然而,现有解释常过于技术化,难以被非专家理解。为此,本文提出一种新流程,利用大、小语言模型构建反事实解释的自然语言叙事。通过知识蒸馏与精炼机制,使小型语言模型在保持强推理能力的同时,表现媲美大型模型。此外,设计了一种简单但有效的评估方法,用于检验生成叙事是否与事实及反事实真值一致。实验表明,该流程显著提升了学生模型的推理能力与实际表现,更适用于真实场景。
原文摘要 · Abstract (English)
Explainable Artificial Intelligence has become a crucial area of research, aiming to demystify the decision-making processes of deep learning models. Among various explainability techniques, counterfactual explanations have been proven particularly promising, as they offer insights into model behavior by highlighting minimal changes that would alter a prediction. Despite their potential, these explanations are often complex and technical, making them difficult for non-experts to interpret. To address this challenge, we propose a novel pipeline that leverages Language Models, large and small, to compose narratives for counterfactual explanations. We employ knowledge distillation techniques along with a refining mechanism to enable Small Language Models to perform comparably to their larger counterparts while maintaining robust reasoning abilities. In addition, we introduce a simple but effective evaluation method to assess natural language narratives, designed to verify whether the models' responses are in line with the factual, counterfactual ground truth. As a result, our proposed pipeline enhances both the reasoning capabilities and practical performance of student models, making them more suitable for real-world use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。