语言模型能学统计规律,但只有GPT-4能分清本质属性与统计规律。
Failures and Successes to Learn a Core Conceptual Distinction from the Statistics of Language
- 用语言数据训练模型,测试其区分本质属性与统计规律的能力。
- 多数模型只看出现频率,无法独立判断属性是否本质,唯有GPT-4做到。
- 对理解语言如何塑造概念、构建因果认知有重要启示。
像“老虎有条纹”这类陈述反映的是本质属性,而“汽车有收音机”仅是统计规律。人类能敏锐区分这两者。有人认为这种区分能力是先天概念结构的一部分,无法通过学习获得。本文研究:能否仅从语言本身习得这一核心概念区分?结果发现,所有语言模型都对统计频率敏感,但在控制频率影响后,均难以识别本质性与统计性之别,唯独GPT-4成功实现。这表明语言经验可能可引导学习深层概念区分,并支持直接从语言构建复杂因果模型的可能性。
原文摘要 · Abstract (English)
Generic statements like "tigers are striped" and "cars have radios" communicate information that is, in general, true. However, while the first statement is true in principle, the second is true only statistically. People are exquisitely sensitive to this principled-vs-statistical distinction. It has been argued that this ability to distinguish between something being true by virtue of it being a category member versus being true because of mere statistical regularity, is a general property of people's conceptual machinery and cannot itself be learned. We investigate whether the distinction between principled and statistical properties can be learned from language itself. If so, it raises the possibility that language experience can bootstrap core conceptual distinctions and that it is possible to learn sophisticated causal models directly from language. We find that language models are all sensitive to statistical prevalence, but struggle with representing the principled-vs-statistical distinction controlling for prevalence. Until GPT-4, which succeeds.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。