arXiv:2607.17625cs.CVcs.LG2026-07

用生物视觉机制提升视频模型效率与可解释性

Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection

论文配图:Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection
图 1 · 摘自论文原文
  • 设计稀疏竞争选票模块替代密集注意力,模拟生物视觉路由
  • 双路径架构:高分辨率'什么'流+低分辨率'哪'流,融合后分类
  • 在脑电数据上相关性达0.18,接近人类感知上限

现代视频变压器通常忽略灵长类视觉原理,且很少基于神经数据评估,限制了其生物可解释性。我们引入一种稀疏胜者通吃令牌选择模块,取代密集自注意力,以提高效率并近似生物视觉回路中的竞争性路由。我们还提出一种受神经启发的分-融视频变压器,采用两条互补路径:高分辨率、低帧率的'什么'流和低分辨率、高帧率的'哪'流,在分类前融合。在Kinetics-400和Something-Something V2数据集上,我们的最优变体在相近规模和预训练条件下,达到准确率与推理时间的帕累托前沿,并对空间扰动表现出更强鲁棒性。通过模型嵌入与相同视频刺激下时间解析脑电记录的表示相似性分析,模型达到峰值脑-模型相关性0.18(约噪声上限的78%),持续优于强视频变压器基线,表明路径特化和稀疏竞争是实现高效、脑对齐视频理解的有效归纳偏置。

原文摘要 · Abstract (English)

Modern video transformers typically ignore principles from primate vision and are rarely evaluated against neural data, limiting their biological interpretability. We introduce a sparse winner-takes-all token selection module that replaces dense self-attention to improve efficiency and approximate competitive routing observed in biological visual circuits. We further propose a neuro-inspired split-and-fuse video transformer which uses two complementary pathways: a high-resolution, low-frame-rate "what" stream and a low-resolution, high-frame-rate "where" stream, fused before classification. On Kinetics-400 and Something-Something V2, our best variant operates on the Pareto frontier of accuracy versus inference time among models of comparable scale and pretraining, and showing improved robustness to spatial perturbations. Using representational similarity analysis between model embeddings and time-resolved EEG recordings for the same video stimuli, our model attains a peak brain-model correlation of 0.18 (about 78% of the noise ceiling) and consistently outperforms strong video transformer baselines, suggesting that pathway specialization and sparse competition are useful inductive biases for efficient, brain-aligned video understanding.

视频理解脑对齐稀疏注意力双路径架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。