arXiv:2504.08001cs.CL2025-04综述被引 15

系统梳理160篇论文,揭示Transformer模型的语义理解能力

Linguistic Interpretability of Transformer-based Language Models: a systematic review

  • 从句法、形态、词汇语义等角度分析模型内部表示
  • 覆盖多语言模型,涵盖预训练阶段而非下游任务
  • 为理解大模型如何学习人类语言提供系统性参考

基于Transformer架构的语言模型在文本分类、情感分析等任务中表现优异,但其内部计算机制仍不清晰,被视为‘黑箱’。近年来,‘可解释性’研究致力于揭示模型如何编码信息,特别是其是否具备类似人类的语言知识,这一领域被称为‘语言可解释性’。本综述系统分析了160篇相关研究,涵盖多种语言与模型(包括多语言模型),从句法、形态、词汇语义和话语等传统语言学视角出发,探索预训练语言模型内部表征中的语言知识。该工作填补了现有可解释性研究的空白,突破了仅关注英文模型或特定下游任务的局限,强调使用内部表示分析技术的研究方法。

原文摘要 · Abstract (English)

Language models based on the Transformer architecture achieve excellent results in many language-related tasks, such as text classification or sentiment analysis. However, despite the architecture of these models being well-defined, little is known about how their internal computations help them achieve their results. This renders these models, as of today, a type of 'black box' systems. There is, however, a line of research -- 'interpretability' -- aiming to learn how information is encoded inside these models. More specifically, there is work dedicated to studying whether Transformer-based models possess knowledge of linguistic phenomena similar to human speakers -- an area we call 'linguistic interpretability' of these models. In this survey we present a comprehensive analysis of 160 research works, spread across multiple languages and models -- including multilingual ones -- that attempt to discover linguistic information from the perspective of several traditional Linguistics disciplines: Syntax, Morphology, Lexico-Semantics and Discourse. Our survey fills a gap in the existing interpretability literature, which either not focus on linguistic knowledge in these models or present some limitations -- e.g. only studying English-based models. Our survey also focuses on Pre-trained Language Models not further specialized for a downstream task, with an emphasis on works that use interpretability techniques that explore models' internal representations.

语言模型可解释性Transformer多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。