神经网络通过眼球运动模拟,自发形成位置与物体绑定的动态表征。
Path Integration and Object-Location Binding Emerge in an Action-Conditioned Predictive Sequence Network
- 用序列采样+眼球运动式位移训练循环网络,模拟认知中的预测过程。
- 在新场景中预测准确率随序列提升,表明上下文学习能力。
- 可后期学习新绑定关系,适合研究认知建模与注意力机制的学者。
适应性认知需要对物体及其关系的结构化内部模型。尽管预测性神经网络常被用来学习这类世界模型,但其具体实现方式及如何支持预测仍不明确。本文在最小化的类脑环境中进行研究:一个循环神经网络从二维连续标记场景中逐个采样标记,并基于当前输入和类似眼动的位移来预测下一个标记。在新场景中,预测准确率随序列推进而提升,表明具备上下文学习能力。解码分析揭示了路径整合与标记身份到位置的动态绑定。干预分析显示,新的绑定可在序列后期学习,且可处理分布外的绑定。这些发现表明,依赖灵活绑定的结构化表征能自发出现以支持预测,为认知科学中的序列世界建模提供了机制解释。
原文摘要 · Abstract (English)
Adaptive cognition requires structured internal models of objects and their relations. Predictive neural networks are often proposed to learn such world models, but how these are instantiated and how they support prediction remain unclear. We investigate this in a minimal in-silico setting. A recurrent neural network samples tokens sequentially from 2D continuous token scenes and is trained to predict the upcoming token from the current input and a saccade-like displacement. On novel scenes, prediction accuracy improves across the sequence, indicating in-context learning. Decoding analyses reveal path integration and dynamic binding of token identity to position. Interventional analyses show that new bindings can be learned late in sequence and that out-of-distribution bindings can be learned as well. Together, these findings show how structured representations relying on flexible binding emerge to support prediction, offering a mechanistic account of sequential world modeling relevant to cognitive science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。