arXiv:2502.20344cs.CL2025-02EMNLP被引 14

用稀疏自编码器解析大模型语言机制,揭示其内在知识表示。

LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder

  • 基于稀疏自编码器构建系统性分析框架,覆盖中英文多维度语言特征。
  • 发现大模型内部存在可识别的语言知识表征,跨层与跨语言分布有规律。
  • 适合对模型可解释性、语言认知机制感兴趣的科研人员参考。

大型语言模型(LLMs)在指代消解、隐喻识别与生成等复杂语言任务上表现卓越,但其处理与表征语言知识的内部机制仍不透明。现有研究受限于粒度粗、分析规模小、聚焦面窄。本文提出LinguaLens,一种基于稀疏自编码器(SAEs)的系统化分析框架,提取中英文四维语言特征(形态、句法、语义、语用)。通过反事实方法构建大规模语言特征反事实数据集,用于机制分析。研究发现:大模型内部存在内在语言知识表征,展现跨层与跨语言分布模式,并具备控制输出的潜力。本工作提供了系统性的资源与方法,证明了大模型具备真实语言知识,为未来更可解释、可控的语言建模奠定基础。

原文摘要 · Abstract (English)

Large language models (LLMs) demonstrate exceptional performance on tasks requiring complex linguistic abilities, such as reference disambiguation and metaphor recognition/generation. Although LLMs possess impressive capabilities, their internal mechanisms for processing and representing linguistic knowledge remain largely opaque. Prior research on linguistic mechanisms is limited by coarse granularity, limited analysis scale, and narrow focus. In this study, we propose LinguaLens, a systematic and comprehensive framework for analyzing the linguistic mechanisms of large language models, based on Sparse Auto-Encoders (SAEs). We extract a broad set of Chinese and English linguistic features across four dimensions (morphology, syntax, semantics, and pragmatics). By employing counterfactual methods, we construct a large-scale counterfactual dataset of linguistic features for mechanism analysis. Our findings reveal intrinsic representations of linguistic knowledge in LLMs, uncover patterns of cross-layer and cross-lingual distribution, and demonstrate the potential to control model outputs. This work provides a systematic suite of resources and methods for studying linguistic mechanisms, offers strong evidence that LLMs possess genuine linguistic knowledge, and lays the foundation for more interpretable and controllable language modeling in future research.

语言模型可解释性自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。