提出五种让大模型内部机制可解释的设计方法,提升可信度。
Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures
- 从功能透明到隐层稀疏,五类架构设计实现模型内在可解释
- 相比事后解释,该方法在结构上直接支持透明性,避免外部近似误差
- 适合关注AI可信、安全部署的研究者和工程师
尽管大语言模型在众多自然语言任务中表现优异,但其内部机制的不透明性限制了可信度与安全部署。现有可解释AI综述多聚焦于事后解释方法,即通过外部近似来解析训练好的模型。与此不同,内在可解释性通过在模型架构与计算过程中直接嵌入透明性,成为新兴且有前景的替代方案。本文系统梳理了大语言模型内在可解释性的最新进展,将现有方法归纳为五种设计范式:功能透明、概念对齐、表征可分解性、显式模块化与潜在稀疏性诱导。我们进一步讨论了当前挑战,并展望该领域的未来研究方向。论文列表见:https://github.com/PKU-PILLAR-Group/Survey-Intrinsic-Interpretability-of-LLMs。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and safe deployment. Existing surveys in explainable AI largely focus on post-hoc explanation methods that interpret trained models through external approximations. In contrast, intrinsic interpretability, which builds transparency directly into model architectures and computations, has recently emerged as a promising alternative. This paper presents a systematic review of the recent advances in intrinsic interpretability for LLMs, categorizing existing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction. We further discuss open challenges and outline future research directions in this emerging field. The paper list is available at: https://github.com/PKU-PILLAR-Group/Survey-Intrinsic-Interpretability-of-LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。