arXiv:2605.24528cs.AIcs.CL2026-05

比较儿童与大模型在不确定性下的推理能力,发现两者行为相似但信息获取方式不同。

Hypothesis Generation and Inductive Inference in Children and Language Models

论文配图:Hypothesis Generation and Inductive Inference in Children and Language Models
图 1 · 摘自论文原文
  • 用程序归纳和贝叶斯粒子推断建模推理过程
  • 大模型复现了儿童对证据可靠性的敏感反应
  • 适合研究认知推理机制或大模型认知类比的学者

现实决策需在证据、潜在因果规则及世界状态均不确定的情况下构建心理模型。人类在该条件下推理的计算原理是什么?基于大模型的智能体在相同约束下是否表现出类似行为?我们通过一个诱导推理盒子任务,让儿童和大模型智能体通过与不确定环境的序列交互来推断隐藏原因。将任务形式化为基于贝叶斯粒子的程序归纳,支持两种互补解释:(1) 假设约束满足过程,(2) 可执行程序的程序合成问题。采用约束框架发现,儿童行为最符合主观证据可靠性与在线假设生成的结合,解释了其证据寻求模式及任务完成与规则泛化之间的分离。使用程序合成框架,将大模型视为可操控的实验系统。在不同后端中,大模型智能体复制了儿童对证据可靠性与可观测性变化的反应,包括忽略不可靠证据、主动寻求信息以解决部分信息困境,以及任务完成与因果泛化间的分离。同时,大模型倾向于过度观察和过度遵守指令。结果表明,尽管儿童与大模型智能体对环境结构的适应性相似,但其信息寻求行为反映出不同的内在成本与归纳偏见。

原文摘要 · Abstract (English)

Real world decision-making requires constructing mental models under uncertainty over evidence, over the underlying causal rules, and over the state of the world itself. Which computational principles underpin human inference under such conditions, and do LLM-based agents exhibit similar behavior given matching constraints? We address these questions using an inductive inference Box Task in which participants, human children and LLM-based agents, infer a latent cause through sequential interaction with an uncertain environment. We formalize this task as program induction with Bayesian particle-based inference, admitting two complementary interpretations: (1) as a constraint satisfaction process over hypotheses, and (2) as a program synthesis problem in which hypotheses are executable programs evaluated against evidence. Using the constraint-based formulation, we show that children's behavior is best explained by a combination of subjective evidence reliability and online hypothesis generation, accounting for both their evidence-seeking patterns and their dissociation between task completion and rule generalization. Using the program synthesis formulation, we treat LLM-based agents as model organisms: controllable systems that allow systematic manipulation of task conditions. Across backends, LLM-based agents replicate children's responses to changes in evidence reliability and observability, including discounting unreliable evidence, seeking to resolve partial information, and dissociating between task completion and causal generalization. At the same time, LLM-based agents tend to over-observe and over-comply with instructions relative to children. These results suggest that while children and LLM-based agents adapt similarly to environmental structure, their information-seeking behavior exhibits distinct underlying costs and inductive biases.

认知推理大模型儿童发展贝叶斯推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。