用循环视觉变压器模拟灵长类视觉注意,表现与生物一致。
A recurrent vision transformer shows signatures of primate visual attention
- 将自注意力与循环记忆结合,让当前输入和存储信息共同引导注意。
- 在空间提示任务中,提示有效性越高,识别准确率提升越明显,反应更快。
- 模型注意图显示动态空间优先机制,适合研究生物注意或神经科学交叉领域。
注意力在生物智能与人工智能中均至关重要,但动物注意与AI自注意力研究长期脱节。我们提出一种循环视觉变压器(Recurrent ViT),将自注意力与循环记忆结合,使当前输入和存储信息共同指导注意分配。该模型仅通过稀疏奖励反馈,在空间提示定向变化检测任务(灵长类研究常用范式)上训练,便展现出灵长类注意的典型特征:提示刺激的识别准确率提升,反应速度加快,且效果随提示有效性增强而上升。自注意力图分析揭示了动态空间优先化,且在预期变化前出现重新激活现象;针对性扰动实验导致性能变化,与灵长类额叶眼区和上丘观察到的神经响应相似。结果表明,将循环反馈引入自注意力可捕捉灵长类视觉注意的关键特征。
原文摘要 · Abstract (English)
Attention is fundamental to both biological and artificial intelligence, yet research on animal attention and AI self attention remains largely disconnected. We propose a Recurrent Vision Transformer (Recurrent ViT) that integrates self-attention with recurrent memory, allowing both current inputs and stored information to guide attention allocation. Trained solely via sparse reward feedback on a spatially cued orientation change detection task, a paradigm used in primate studies, our model exhibits primate like signatures of attention, including improved accuracy and faster responses for cued stimuli that scale with cue validity. Analysis of self-attention maps reveals dynamic spatial prioritization with reactivation prior to expected changes, and targeted perturbations produce performance shifts similar to those observed in primate frontal eye fields and superior colliculus. These findings demonstrate that incorporating recurrent feedback into self attention can capture key aspects of primate visual attention.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。