语言模型无法仅靠文本完全理解语义,存在根本性信息瓶颈。
A Formal Limitation on Learning Human Language From Textual Corpora

- 从信息论角度证明:仅凭文本无法完全还原说话人意图。
- 理论上限由语言固有不确定性决定,与模型大小无关。
- 适用于所有基于文本的模型,包括大语言模型,适合理论研究者。
能否仅从话语形式中推断说话人的真实意图?我们从信息论角度回答此问题,针对任意文本特征提取器(包括当代大语言模型的隐藏状态)。将语言使用建模为意义、语境与话语的联合分布,推导出解码器从话语表示中恢复说话人意图的概率上界。该上界由形式对意义的不确定性决定,其包含不可消除的部分以及仅依赖外部语境(而非话语本身)才能解决的部分。由于这些量是语言的本质属性,任何表示(无论训练数据多少或监督程度如何)都无法超越此界限。该结论在离散或连续意义空间下均成立。人工语言、中文零代词消解及颜色指称任务的实验提供了理论支持。
原文摘要 · Abstract (English)
Can a listener recover what a speaker means from the form of an utterance alone? We answer this question information-theoretically, and for a listener given by any featurizer of text, including the hidden states of contemporary large language models. Modeling language use as a joint distribution over meanings, contexts, and utterances, we derive upper bounds on the probability that a decoder recovers a speaker's intended meaning from a representation of the utterance. The bounds are governed by the uncertainty that form leaves about meaning, which splits into an irreducible part and a part that only (extralinguistic) context, but never the utterance alone, can resolve. Because these quantities are intrinsic to language, no representation, however much text or supervision produced it, can surpass them; the bounds hold whether the space of meanings is discrete or continuous. Experiments on artificial languages, Mandarin zero-pronoun resolution, and color reference provide empirical evidence in support of the theory.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。