用能量模型模拟视觉语言融合认知,解释注意力错觉与语言理解现象。
A Hierarchical Energy-Based Model for Multimodal Cognition
- 构建分层能量模型,让视觉与语言预测共享一个无模态中枢。
- 能解释忽视盲视、翻转立方体等注意力现象,还原语言认知的脑电特征。
- 适合研究认知机制、神经科学建模或对大模型提出新对比思路的人。
我们提出IM-LEPP(整合多模态潜在能量预测处理),一种分层的能量基多模态认知模型,将先前单模态模型LEPP扩展至融合视觉与语言。基于生成神经网络是认知动力学的有效理论这一观点,类似统计力学与热力学的关系,IM-LEPP将认知建模为潜在状态在学习到的能量景观中流动,而非神经回路的直接描述。其架构采用枢纽-辐条结构,依据Lambon Ralph等人的受控语义认知框架,视觉对象、场景与语言单元的预测编码路径汇聚于一个共享的无模态中枢(模拟前颞叶)。每条路径的预测由当前中枢状态条件决定,而非被覆盖,从而保持路径特异性的同时反映全模态上下文。该模型可机械性解释注意现象如忽视盲视、奈克立方体双稳态,并恢复或启发心理学中独立验证的发现,包括预期意外理论、N400/P600脑电成分、花园路径重分析;同时提出与大语言模型在词语预测轨迹敏感性上的可检验对比。还讨论了数据高效语言习得、语义/情景记忆子系统设计,以及与预测编码、自由能原理、JEPA、层次时间记忆的定位关系,并提出具体实验预测以验证核心主张。
原文摘要 · Abstract (English)
We propose IM-LEPP (Integrated Multimodal Latent Energy-based Predictive Processing), a hierarchical, energy-based model of multimodal cognition that extends a previously proposed single-modality model (LEPP) to integrate vision and language. Following the view that generative neural networks are effective theories of cognitive dynamics, analogous to how statistical mechanics relates to thermodynamics, IM-LEPP models cognition as latent states flowing through learned energy landscapes rather than as an account of neural circuitry. The architecture is a hub-and-spoke hierarchy, grounded in the controlled semantic cognition framework of Lambon Ralph et al., in which predictive-coding pipelines for visual objects, scenes, and linguistic units converge on a shared amodal hub modeled on the anterior temporal lobe. Each pipeline's own prediction is conditioned by, rather than overwritten by, the current hub state, preserving pipeline-specific identity while letting every prediction reflect the full multimodal context. We show this architecture gives a mechanistic account of attentional phenomena such as inattentional blindness and Necker-cube bistability, and that its structure recovers or motivates independently established findings in psycholinguistics, including surprisal theory, the N400/P600 ERP components, and garden-path reanalysis, alongside a falsifiable contrast with transformer language models on trajectory-sensitivity in next-word prediction. We also discuss data-efficient language acquisition relative to LLMs, outline a semantic/episodic memory subsystem, situate the model against predictive coding, the free-energy principle, JEPA, and Hierarchical Temporal Memory, and propose concrete experimental predictions to test its central claims Key Words: predictive processing; predictive coding; energy-based models; diffusion models; effective theory; computational neuroscience.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。