arXiv:2506.21812cs.CLcs.CV2025-06综述被引 14

梳理大模型可解释性方法,助其从黑箱变透明

Towards Transparent AI: A Survey on Explainable Large Language Models

  • 按Transformer架构分类大模型解释技术:编码器、解码器、混合型
  • 系统评估解释效果,分析实际应用中的可解释性表现
  • 适合关注AI伦理与可信部署的研究者和从业者

大型语言模型(LLMs)在人工智能发展中起关键作用,但其决策过程难以解释,呈现‘黑箱’特性,制约其在高风险场景的应用。为解决这一问题,研究者提出了多种可解释人工智能(XAI)方法以提供人类可理解的解释。然而,对这些方法的系统性认知仍不足。本综述基于LLM的Transformer架构(编码器仅用型、解码器仅用型、编码器-解码器型)对XAI方法进行分类,评估其解释能力,并探讨解释在实际应用中的使用方式。最后,总结现有资源、现存挑战与未来方向,旨在推动更透明、负责任的LLM发展。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have played a pivotal role in advancing Artificial Intelligence (AI). However, despite their achievements, LLMs often struggle to explain their decision-making processes, making them a 'black box' and presenting a substantial challenge to explainability. This lack of transparency poses a significant obstacle to the adoption of LLMs in high-stakes domain applications, where interpretability is particularly essential. To overcome these limitations, researchers have developed various explainable artificial intelligence (XAI) methods that provide human-interpretable explanations for LLMs. However, a systematic understanding of these methods remains limited. To address this gap, this survey provides a comprehensive review of explainability techniques by categorizing XAI methods based on the underlying transformer architectures of LLMs: encoder-only, decoder-only, and encoder-decoder models. Then these techniques are examined in terms of their evaluation for assessing explainability, and the survey further explores how these explanations are leveraged in practical applications. Finally, it discusses available resources, ongoing research challenges, and future directions, aiming to guide continued efforts toward developing transparent and responsible LLMs.

可解释AI大模型透明性XAI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。