通过最小语义对分析,揭示语言模型概率与语法正确性的关系。
What Can String Probability Tell Us About Grammaticality?
- 基于语料生成假设,建立语法、语义与字符串概率的理论框架。
- 在英汉28万组最小语义对中验证三类相关性预测。
- 发现语法错误句与正确句在概率空间中难以区分,提示评估新方向。
语言模型(LM)究竟学到了多少语法知识?这一问题在语言学界仍有争议,但概率与语法正确性在语言学中是两个不同概念,因此字符串概率能否反映模型的语法知识尚不明确。本文基于语料生成过程的简单假设,提出一个关于语法、语义与字符串概率之间关系的理论分析框架,并验证了三个预测:(1)在最小语义对中,字符串概率存在相关性;(2)模型与人类对最小语义对的判断差异(delta)呈正相关;(3)语法正确与错误句子在概率空间中分离效果差。基于英汉共28万组句子对的实证分析,为利用概率研究语言模型结构知识提供了理论依据,并指明未来评估语言模型语法能力的研究方向。
原文摘要 · Abstract (English)
What have language models (LMs) learned about grammar? This question remains hotly debated, with major ramifications for linguistic theory. However, since probability and grammaticality are distinct notions in linguistics, it is not obvious what string probabilities can reveal about an LM's underlying grammatical knowledge. We present a theoretical analysis of the relationship between grammar, meaning, and string probability, based on simple assumptions about the generative process of corpus data. Our framework makes three predictions, which we validate empirically using 280K sentence pairs in English and Chinese: (1) correlation between the probability of strings within minimal pairs, i.e., string pairs with minimal semantic differences; (2) correlation between models' and humans' deltas within minimal pairs; and (3) poor separation in probability space between unpaired grammatical and ungrammatical strings. Our analyses give theoretical grounding for using probability to learn about LMs' structural knowledge, and suggest directions for future work in LM grammatical evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。