发现语言模型内部表征能清晰区分语法正确与错误句子。
Linear representations of grammaticality in neural language models

- 用质心探测法检验模型表征空间中语法句与非语法句的分离程度。
- 语法差异在多种模型、语法现象和语言中均存在显著表征分离。
- 该表征独立于词汇频率等混淆因素,适合评估模型语法能力。
神经语言模型是否具备基于语法正确性区分句子的能力,一直是计算语言学中的争议话题。现有研究多依赖概率指标,但该方法因假设语法正确性与概率天然纠缠而受质疑——概率受词频、语义合理性及世界知识等多种因素影响。本文突破概率评价范式,考察语法正确性是否编码于模型内部表征中。采用质心探测法(mass-mean probing),检验语法句与非语法句在表征空间中是否存在系统性分离。进一步分析该表征是否独立于与语法相关的其他句法属性,并检验其在不同语法现象和语言间的泛化能力。结果表明,多种预训练语言模型的句子表征中均存在稳健的语法正确性编码,可实现显著的表征分离,且无法被其他句级因素完全解释。该编码在广泛语法现象中具有泛化性,并部分跨语言成立,说明语法正确性构成当代语言模型中一个连贯的表征维度。研究为语言模型句法知识的本质提供了新证据,提出了一种不依赖字符串概率的语法能力评估互补框架。
原文摘要 · Abstract (English)
Whether neural language models (NLMs) possess the ability to distinguish strings on the basis of their grammaticality remains a debated topic in the computational linguistics literature. Existing evidence has largely relied on probability-based measures, testing whether models assign higher probabilities to grammatical than ungrammatical strings. However, probability comparisons have been criticized as a measure for grammatical knowledge based on the assumption that grammaticality is inherently entangled with likelihood. Model-assigned probability is a function of many related sentence properties, such as lexical frequency, plausibility, and world knowledge. In this work, we move beyond probability-based evaluations and investigate whether grammaticality is encoded in the internal representations of NLMs. Using mass-mean probing, we test whether grammatical and ungrammatical sentences are systematically separated in representational space. We further examine the extent to which these representations are independent of sentence properties that are correlated with grammaticality, as well as their generalization across grammatical phenomena and languages. Our results provide evidence that grammaticality is robustly encoded in sentence representations of a wide range of pretrained NLMs, yielding clear representational separation on the dimension of grammaticality that cannot be fully explained by alternative sentence-level factors. Moreover, this encoding generalizes across a broad range of grammatical phenomena and to some degree, across languages, suggesting that grammaticality constitutes a coherent representational dimension in contemporary NLMs. These findings contribute new evidence to debates about the nature of syntactic knowledge in language models and offer a complementary framework for evaluating grammatical competence that is not dependent on string probabilities alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。