用游戏脑电数据验证视觉语言与动作模型的神经对齐,发现动作模型更贴近人类大脑活动。
Brain alignment of reasoning and action representations from vision-language and action models during naturalistic gameplay

- 对比视觉语言模型和动作模型在游戏中的神经表征对齐效果。
- 动作模型在前额-顶叶和运动规划区对齐度显著更高,且优于强化学习基线。
- 动作模型的表征组织呈不对称性,尤其在前运动皮层中更贴近真实大脑计算模式。
理解人类与人工智能如何通过环境交互进行预测与规划,是神经科学与机器学习交叉领域的根本挑战。现有脑编码研究多聚焦于语言理解或被动视觉处理,而交互式脑对齐研究主要限于强化学习代理与理论模型。本文利用玩家在自然主义雅达利风格游戏中的fMRI数据,考察两类基础模型——视觉语言模型(VLMs)与大动作模型(LAMs)——的内部表示如何与大脑活动对齐。重点分析以动作导向与推理导向提示引导下模型表征的变化。结果表明:1)两者在体素级编码性能上均显著优于强化学习基线,且该优势在特征维度匹配条件下依然成立;2)提示驱动的提升随皮层加工层级递增:前额-顶叶与运动规划区获益最大,早期视觉皮层仅获得约一半提升;3)方差分解揭示质性差异:VLM呈现提示对称性(12.5%独特动作 vs. 13.6%独特推理),而LAM为提示非对称性(27%独特动作 vs. -5%独特推理),其不对称性在前运动皮层最强。综上,即使整体全脑预测精度相当,动作专用微调仍可使多模态表示向动作相关神经计算重组。
原文摘要 · Abstract (English)
Understanding how humans and artificial intelligence systems predict and plan by interacting with their environment is a fundamental challenge at the intersection of neuroscience and machine learning. Most brain-encoding studies focus on aligning artificial models with brain activity during language comprehension or passive visual processing, while interactive brain-alignment studies have to date been largely limited to reinforcement-learning (RL) agents and theory-based models. To address this gap, we study brain alignment of representative models from two foundation-model families, namely vision-language models (VLMs) and large-action models (LAMs), using fMRI recordings from participants playing naturalistic Atari-style video games. Specifically, we examine how action-focused and reasoning-focused prompts shape model's internal representations and align with fMRI brain activity. First, we find that both VLMs and LAMs exhibit significantly exhibit voxel-wise encoding performance than RL baselines, with the advantage holding even under matched feature dimensionality. Second, prompt-driven gains scale with the cortical processing hierarchy: the largest improvements appear in frontal-parietal and motor-planning regions, while early visual cortex gains roughly half as much. Third, variance partitioning reveals a qualitatively different representational organization: VLM is prompt-symmetric (12.5% unique action vs. 13.6% unique reasoning), whereas LAM is prompt-asymmetric (27% unique action vs. -5% unique reasoning), with the asymmetry strongest in frontal-motor cortex. Together, these results demonstrate that action-specialized fine-tuning reorganizes multimodal representations toward action-relevant neural computations even when whole-brain prediction accuracy is statistically equivalent between VLM and LAM.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。