语言模型能认出物体形状,却无法理解'形状定义类别'这一深层规律。
Exemplar Retrieval Without Overhypothesis Induction: Limits of Distributional Sequence Learning in Early Word Learning
- 用自回归变换器在合成语料上训练,模拟儿童学习过程
- 模型第一层识别准确率100%,第二层泛化能力仅50-52%(随机水平)
- 适合研究儿童认知发展与神经网络归纳偏置的对比
儿童不仅学会球是圆的、积木是方的,更懂得‘形状’是定义物体类别的关键特征——这种关于特征层级的抽象称为超假设。什么学习机制足以实现这种归纳飞跃?我们使用参数量340万至2560万的自回归变压器模型,在包含八种控制条件的合成语料上进行训练,其中形状是跨类别的稳定特征。在预注册的120次运行中,所有模型在1,040项wug测试中均达到100%的第一层实例检索准确率,但对新名词的第二层泛化能力仅为50-52%(接近随机),经等效性检验确认。特征交换诊断显示,模型依赖框架到特征的模板匹配,而非结构化的词—域—特征抽象。结果揭示了在发展规模训练条件下,自回归分布序列学习存在明确局限。
原文摘要 · Abstract (English)
Background: Children do not simply learn that balls are round and blocks are square. They learn that shape is the kind of feature that tends to define object categories -- a second-order generalisation known as an overhypothesis [1, 2]. What kind of learning mechanism is sufficient for this inductive leap? Methods: We trained autoregressive transformer language models (3.4M-25.6M parameters) on synthetic corpora in which shape is the stable feature dimension across categories, with eight conditions controlling for alternative explanations. Results: Across 120 pre-registered runs evaluated on a 1,040-item wug test battery, every model achieved perfect first-order exemplar retrieval (100%) while second-order generalisation to novel nouns remained at chance (50-52%), a result confirmed by equivalence testing. A feature-swap diagnostic revealed that models rely on frame-to-feature template matching rather than structured noun-to-domain-to-feature abstraction. Conclusions: These results reveal a clear limitation of autoregressive distributional sequence learning under developmental-scale training conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。