arXiv:2506.05136cs.CL2025-06ACL被引 11

发现语言局部统计结构影响模型学习难易,揭示神经语言模型的归纳偏置。

Information Locality as an Inductive Bias for Neural Language Models

  • 用局部熵衡量语言的局部不确定性,量化前文对下一个词的消歧能力。
  • 局部熵越高,Transformer和LSTM在自然语料与概率有限状态自动机上的学习越困难。
  • 为理解模型与人类语言处理的共性提供了可量化的信息论框架,适合关注模型归纳偏置的研究者。

归纳偏倚是每个机器学习系统固有的特性,决定了模型如何从有限数据中泛化。对于神经语言模型(LMs),其归纳偏倚是否与人类处理约束一致仍存在争议。为此,我们提出一个定量框架,可对这些偏倚进行受控研究。在该框架中,我们引入了m-局部熵——一种基于平均有损上下文意外性的信息论度量,用于刻画语言的局部不确定性,即前m-1个符号有效消歧下一个符号的程度。在扰动的自然语言语料库和由概率有限状态自动机(PFSAs)定义的语言上进行实验,结果表明:局部熵越高的语言,Transformer和LSTM语言模型越难以学习。这表明神经语言模型与人类一样,对语言的局部统计结构高度敏感。

原文摘要 · Abstract (English)

Inductive biases are inherent in every machine learning system, shaping how models generalize from finite data. In the case of neural language models (LMs), debates persist as to whether these biases align with or diverge from human processing constraints. To address this issue, we propose a quantitative framework that allows for controlled investigations into the nature of these biases. Within our framework, we introduce $m$-local entropy$\unicode{x2013}$an information-theoretic measure derived from average lossy-context surprisal$\unicode{x2013}$that captures the local uncertainty of a language by quantifying how effectively the $m-1$ preceding symbols disambiguate the next symbol. In experiments on both perturbed natural language corpora and languages defined by probabilistic finite-state automata (PFSAs), we show that languages with higher $m$-local entropy are more difficult for Transformer and LSTM LMs to learn. These results suggest that neural LMs, much like humans, are highly sensitive to the local statistical structure of a language.

语言模型归纳偏置信息论局部熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。