arXiv:2604.12543cs.AI2026-04中稿 · publication at the…被引 2

用两阶段大模型框架让AI解释更可信、易懂。

A Two-Stage LLM Framework for Accessible and Verified XAI Explanations

论文配图:A Two-Stage LLM Framework for Accessible and Verified XAI Explanations
图 1 · 摘自论文原文
  • 先生成解释,再由另一模型验证其准确性和完整性。
  • 验证后解释的可读性提升,错误率显著下降。
  • 适合需要可靠AI解释的医疗、金融等高风险领域。

大型语言模型(LLM)正被用于将可解释人工智能(XAI)技术的技术输出转化为自然语言解释。然而,现有方法缺乏准确性、忠实性和完整性保障。当前评估也多依赖主观评分,无法阻止错误解释传递给用户。为此,本文提出一种两阶段大模型元验证框架:首先由解释器大模型将原始XAI输出转化为自然语言叙述;其次由验证器大模型从忠实性、连贯性、完整性及幻觉风险等方面进行评估,并通过迭代反馈机制优化解释内容。在五个XAI技术与数据集上,使用三类开源权重大模型的实验表明,验证能有效过滤不可靠解释,同时提升语言可读性。对优化过程中熵产生率(EPR)的分析显示,验证反馈逐步引导解释器趋向更稳定、一致的推理路径。整体而言,该框架为构建更可信、更普惠的XAI系统提供了高效路径。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly used to translate the technical outputs of eXplainable Artificial Intelligence (XAI) methods into accessible natural-language explanations. However, existing approaches often lack guarantees of accuracy, faithfulness, and completeness. At the same time, current efforts to evaluate such narratives remain largely subjective or confined to post-hoc scoring, offering no safeguards to prevent flawed explanations from reaching end-users. To address these limitations, this paper proposes a Two-Stage LLM Meta-Verification Framework that consists of (i) an Explainer LLM that converts raw XAI outputs into natural-language narratives, (ii) a Verifier LLM that assesses them in terms of faithfulness, coherence, completeness, and hallucination risk, and (iii) an iterative refeed mechanism that uses the Verifier's feedback to refine and improve them. Experiments across five XAI techniques and datasets, using three families of open-weight LLMs, show that verification is crucial for filtering unreliable explanations while improving linguistic accessibility compared with raw XAI outputs. In addition, the analysis of the Entropy Production Rate (EPR) during the refinement process indicates that the Verifier's feedback progressively guides the Explainer toward more stable and coherent reasoning. Overall, the proposed framework provides an efficient pathway toward more trustworthy and democratized XAI systems.

可解释AI大模型验证机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。