语言类型影响大模型的归纳推理能力,与人类儿童相似。
Bears, all bears, and some bears. Language Constraints on Language Models' Inductive Inferences
- 用图像+语言任务测试大模型对'所有''有些''泛指'语句的推理差异。
- 模型表现与4岁以上儿童一致:'所有' > '泛指' > '有些'的推理强度。
- 差异源于归纳约束而非表面形式,适合研究认知建模的学者参考。
语言对归纳推理施加微妙限制。Gelman等(2002)发现,4岁及以上儿童在将新属性推广到特定对象时,能区分泛指句('熊是daxable')、全称句('所有熊都是daxable')和不定复数句('有些熊是daxable'),推理强度依次递减(所有 > 泛指 > 有些),表明他们对这些命题有不同表征。我们通过复现原实验,测试通用统计学习者如视觉-语言模型是否也表现出类似差异。在一系列前提测试(图像中稳健识别类别、对'all'和'some'敏感)后,模型在推理任务中表现出与人类行为一致的模式。事后分析其表示层发现,这种差异基于归纳约束而非表面形式差异。
原文摘要 · Abstract (English)
Language places subtle constraints on how we make inductive inferences. Developmental evidence by Gelman et al. (2002) has shown children (4 years and older) to differentiate among generic statements ("Bears are daxable"), universally quantified NPs ("all bears are daxable") and indefinite plural NPs ("some bears are daxable") in extending novel properties to a specific member (all > generics > some), suggesting that they represent these types of propositions differently. We test if these subtle differences arise in general purpose statistical learners like Vision Language Models, by replicating the original experiment. On tasking them through a series of precondition tests (robust identification of categories in images and sensitivities to all and some), followed by the original experiment, we find behavioral alignment between models and humans. Post-hoc analyses on their representations revealed that these differences are organized based on inductive constraints and not surface-form differences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。