arXiv:2512.13517q-bio.NCcs.LG2025-12中稿 · ICML被引 2

用深度学习模拟人类心理旋转,结合虚拟现实实验验证

A Deep Learning Model of Mental Rotation Informed by Interactive VR Experiments

  • 三阶段模型:等变编码器、符号化编码、神经决策代理
  • 准确复现实验中的人类反应时间与行为模式
  • 适合研究认知建模或空间推理的学者参考

心理旋转——从不同视角比较物体的能力——是人类心智模拟与空间建模的基本范例。本文提出一种机制性的人类心理旋转模型,融合深度、等变及神经符号学习的最新进展。模型由三部分组成:(1) 等变神经编码器,从图像生成物体的3D空间表征;(2) 神经符号物体编码器,从中提取符号化描述;(3) 神经决策代理,通过循环路径在3D隐空间中生成旋转模拟,用于比较符号描述。模型设计基于现有心理旋转实验文献,并通过虚拟现实实验补充数据,参与者可交互操控物体进行对比。模型成功捕捉了我们及他人实验中的表现、反应时间与行为特征,且消融实验证明各组件必要性。本工作丰富了深度神经模型对人类空间推理的建模体系,进一步证明整合深度、等变与符号表征在建模人类心智中的有效性。

原文摘要 · Abstract (English)

Mental rotation -- the ability to compare objects seen from different viewpoints -- is a fundamental example of mental simulation and spatial world modeling in humans. Here we propose a mechanistic model of human mental rotation, leveraging recent advances in deep, equivariant, and neuro-symbolic learning. Our model consists of three stacked components: (1) an equivariant neural encoder, producing 3D spatial representations of objects from images, (2) a neuro-symbolic object encoder, deriving symbolic objects descriptions from these spatial representations, and (3) a neural decision agent, comparing these symbolic descriptions to prescribe rotation simulations in 3D latent space via a recurrent pathway. Our model design is guided by the existing experimental literature on mental rotation, which we complemented with experiments in VR where participants could at times manipulate the objects to compare. Our model captures well the performance, response times and behavior of participants in our and others' experiments, and through ablation studies we demonstrate the necessity of each component. Our work adds to a recent collection of deep neural models of human spatial reasoning, further demonstrating the potency of integrating deep, equivariant, and symbolic representations to model the human mind.

心理旋转神经符号虚拟现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。