AI模型需先在具体环境里学习才能迁移因果知识,而人类能直接抽象应用。
Grounding Before Generalizing: How AI Differs from Humans in Causal Transfer
- 通过序列探索发现共因与共果结构,验证模型是否具备类人因果迁移能力。
- 模型在视觉条件下表现更差,且需先建立环境特异性映射才有效率提升。
- 模型存在共因/共果判断偏差,反映其缺乏人类的无偏因果抽象机制。
提取抽象因果结构并应用于新情境是人类智能的核心特征。尽管大语言模型(LLMs)和视觉语言模型(VLMs)在多种推理任务中表现优异,但它们在交互式因果学习——通过序列探索推断潜在结构并跨情境迁移——方面的能力仍不清楚。人类学习者在极少暴露下即可完成迁移,而经典强化学习(RL)代理则会灾难性失败。当前人工智能模型是否具备类似人类的抽象因果结构迁移机制尚不明确。基于OpenLock范式,该研究要求模型依次发现共因(CC)和共果(CE)结构,结果表明:模型表现出根本性的迁移延迟或缺失——即使成功模型也需先建立环境特异性映射(即‘环境奠基’),效率提升才出现;而人类则从首次解题即利用已有结构知识。在仅文本条件下,模型表现匹配或超过人类;但在图像仅存及图文并存条件下,视觉信息整体降低而非提升性能,揭示模型广泛依赖符号处理而非融合多模态推理。模型还表现出系统性共因/共果不对称性,与人类无关,暗示其存在启发式偏差而非方向中立的因果抽象。这些发现表明,大规模统计学习无法生成支撑人类类比推理的去情境化因果模式,确立‘奠基性迁移’是当前LLMs和VLMs的根本局限。
原文摘要 · Abstract (English)
Extracting abstract causal structures and applying them to novel situations is a hallmark of human intelligence. While Large Language Models (LLMs) and Vision Language Models (VLMs) have shown strong performance on a wide range of reasoning tasks, their capacity for interactive causal learning -- inducing latent structures through sequential exploration and transferring them across contexts -- remains uncharacterized. Human learners accomplish such transfer after minimal exposure, whereas classical Reinforcement Learning (RL) agents fail catastrophically. Whether state-of-the-art Artificial Intelligence (AI) models possess human-like mechanisms for abstract causal structure transfer is an open question. Using the OpenLock paradigm requiring sequential discovery of Common Cause (CC) and Common Effect (CE) structures, here we show that models exhibit fundamentally delayed or absent transfer: even successful models require initial environmental-specific mapping -- what we term environmental grounding -- before efficiency gains emerge, whereas humans leverage prior structural knowledge from the very first solution attempt. In the text-only condition, models matched or exceeded human discovery efficiency. In contrast, visual information -- in both the image-only and text-and-image conditions -- overall degraded rather than enhanced performance, revealing a broad reliance on symbolic processing rather than integrated multimodal reasoning. Models further exhibited systematic CC/CE asymmetries absent in humans, suggesting heuristic biases rather than direction-neutral causal abstraction. These findings reveal that large-scale statistical learning does not produce the decontextualized causal schemas underpinning human analogical reasoning, establishing grounding-dependent transfer as a fundamental limitation of current LLMs and VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。