为婴儿语言模型设计发展心理学驱动的推理评测基准。
BabyReasoningBench: Generating Developmentally-Inspired Reasoning Tasks for Evaluating Baby Language Models
- 基于儿童心理学经典范式生成19个推理任务
- 两个基于GPT-2的婴儿模型在因果推理上表现较好
- 适合研究儿童语言训练下推理能力的形成机制
传统语言模型推理评估多基于成人中心的基准,假设具备广泛世界知识、复杂指令遵循和成熟语用能力,这些假设与以儿童导向语音和早期叙事文本训练的婴儿语言模型不匹配,掩盖了其在约束条件下可能涌现的推理能力。我们提出BabyReasoningBench,一个由GPT-5.2生成的19项推理任务基准,涵盖心理理论、类比与关系推理、因果推断与干预选择,以及受记忆和语用干扰的核心推理原语。实验发现,两个基于GPT-2、分别在10M和100M儿童导向文本上预训练的婴儿模型整体表现较低但分布不均:规模扩大改善了部分因果与物理推理任务,而信念归属和语用敏感任务仍具挑战。该基准为分析儿童式训练数据支持的推理类型提供了发展学基础视角,并可用于检验相关能力形成的机制假说。
原文摘要 · Abstract (English)
Traditional evaluations of reasoning capabilities of language models are dominated by adult-centric benchmarks that presuppose broad world knowledge, complex instruction following, and mature pragmatic competence. These assumptions are mismatched to baby language models trained on developmentally plausible input such as child-directed speech and early-childhood narratives, and they obscure which reasoning abilities (if any) emerge under such constraints. We introduce BabyReasoningBench, a GPT-5.2 generated benchmark of 19 reasoning tasks grounded in classic paradigms from developmental psychology, spanning theory of mind, analogical and relational reasoning, causal inference and intervention selection, and core reasoning primitives that are known to be confounded by memory and pragmatics. We find that two GPT-2 based baby language models (pretrained on 10M and 100M of child-directed speech text) show overall low but uneven performance, with dissociations across task families: scaling improves several causal and physical reasoning tasks, while belief attribution and pragmatics-sensitive tasks remain challenging. BabyReasoningBench provides a developmentally grounded lens for analyzing what kinds of reasoning are supported by child-like training distributions, and for testing mechanistic hypotheses about how such abilities emerge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。