系统梳理大模型可解释性技术,分类解析不同架构的解释方法。
Towards Transparent AI: A Survey on Explainable Language Models
- 按编码器、解码器等结构分类,梳理适配不同大模型的可解释技术
- 从合理性和忠实性双维度评估现有方法的有效性
- 适合关注AI透明性、模型可信度的研究者与应用开发者
语言模型在自然语言处理中取得显著进展,但其黑箱特性引发对内部机制和决策过程可解释性的关切,尤其在高风险领域,利益相关方需理解模型输出依据以确保问责。尽管可解释人工智能(XAI)在非语言模型中已有深入研究,但应用于语言模型时面临复杂架构、海量训练数据及强泛化能力等挑战。现有综述多未充分反映模型架构多样性与能力演进带来的新问题。本文系统回顾了面向语言模型的可解释性技术,依据其底层Transformer架构(编码器-仅、解码器-仅、编码器-解码器)进行组织,分析各类方法如何适配并评估其优劣。同时,通过合理性与忠实性双重标准评估技术效果。最后,识别开放挑战并提出未来方向,旨在推动构建更鲁棒、透明且可解释的可解释性方法。
原文摘要 · Abstract (English)
Language Models (LMs) have significantly advanced natural language processing and enabled remarkable progress across diverse domains, yet their black-box nature raises critical concerns about the interpretability of their internal mechanisms and decision-making processes. This lack of transparency is particularly problematic for adoption in high-stakes domains, where stakeholders need to understand the rationale behind model outputs to ensure accountability. On the other hand, while explainable artificial intelligence (XAI) methods have been well studied for non-LMs, they face many limitations when applied to LMs due to their complex architectures, considerable training corpora, and broad generalization abilities. Although various surveys have examined XAI in the context of LMs, they often fail to capture the distinct challenges arising from the architectural diversity and evolving capabilities of these models. To bridge this gap, this survey presents a comprehensive review of XAI techniques with a particular emphasis on LMs, organizing them according to their underlying transformer architectures: encoder-only, decoder-only, and encoder-decoder, and analyzing how methods are adapted to each while assessing their respective strengths and limitations. Furthermore, we evaluate these techniques through the dual lenses of plausibility and faithfulness, offering a structured perspective on their effectiveness. Finally, we identify open research challenges and outline promising future directions, aiming to guide ongoing efforts toward the development of robust, transparent, and interpretable XAI methods for LMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。