arXiv:2505.18970cs.CL2025-05被引 3

用原型生成可解释的模型,让大模型决策更透明可信。

Learning to Explain: Prototype-Based Surrogate Models for LLM Classification

  • 基于句子原型构建可解释的替代模型,贴近大模型推理逻辑。
  • 在多个数据集上优于现有方法,且少量数据即可高效训练。
  • 适合需要高可信度解释的医疗、金融等实际场景使用。

大型语言模型(LLMs)在自然语言任务中表现优异,但其决策过程仍高度不透明。现有解释方法或忠实度不足,或人类难以理解。为此,我们提出一种新型原型驱动的替代框架 ProtoSurE,通过设计可解释的代理模型,结合句级原型作为人类可理解的概念,实现对目标 LLM 的忠实且易懂解释。大量实验表明,ProtoSurE 在多种 LLM 和数据集上均持续优于当前最优解释方法。尤其具备强数据效率,仅需少量训练样本即可获得良好性能,适用于真实应用场景。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated impressive performance on natural language tasks, but their decision-making processes remain largely opaque. Existing explanation methods either suffer from limited faithfulness to the model's reasoning or produce explanations that humans find difficult to understand. To address these challenges, we propose \textbf{ProtoSurE}, a novel prototype-based surrogate framework that provides faithful and human-understandable explanations for LLMs. ProtoSurE trains an interpretable-by-design surrogate model that aligns with the target LLM while utilizing sentence-level prototypes as human-understandable concepts. Extensive experiments show that ProtoSurE consistently outperforms SOTA explanation methods across diverse LLMs and datasets. Importantly, ProtoSurE demonstrates strong data efficiency, requiring relatively few training examples to achieve good performance, making it practical for real-world applications.

可解释性大模型原型学习代理模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。