现有视频神经网络仍无法模拟猴子大脑对动态物体的不变性感知。
Better, But Not Sufficient: Testing Video ANNs Against Macaque IT Dynamics
- 用自然视频测试猴子脑区响应,对比静态、循环和视频类神经网络
- 视频模型在后期响应阶段预测力提升,但对去形变视频仍失败
- 适合研究生物视觉动态计算与下一代视频模型设计者
基于静态图像训练的前馈人工神经网络(ANN)仍是灵长类腹侧视觉通路的主要模型,但其固有局限在于仅支持静态计算。灵长类世界是动态的,猕猴腹侧视觉通路尤其是下颞叶(IT)皮层不仅支持物体识别,还编码自然视频观看中的物体运动速度。IT的时序响应是否仅反映时间展开的前馈变换、逐帧特征加浅层时序池化,还是蕴含更丰富的动态计算?我们通过比较猕猴IT在自然视频下的响应与静态、循环及视频型ANN模型的表现进行测试。视频模型在后期响应阶段带来适度的神经可预测性提升,引发对其捕捉动态机制的疑问。为此,我们引入压力测试:在自然视频上训练的解码器被评估于‘去形变’版本视频(保留运动但移除形状和纹理)。结果显示,IT种群活动能跨此操作泛化,但所有类型的ANN均失败。这表明当前视频模型更擅长捕捉依赖外观的动态,而非IT中体现的外观不变性时序计算,凸显需建立新目标以编码生物时序统计特性与不变性。
原文摘要 · Abstract (English)
Feedforward artificial neural networks (ANNs) trained on static images remain the dominant models of the the primate ventral visual stream, yet they are intrinsically limited to static computations. The primate world is dynamic, and the macaque ventral visual pathways, specifically the inferior temporal (IT) cortex not only supports object recognition but also encodes object motion velocity during naturalistic video viewing. Does IT's temporal responses reflect nothing more than time-unfolded feedforward transformations, framewise features with shallow temporal pooling, or do they embody richer dynamic computations? We tested this by comparing macaque IT responses during naturalistic videos against static, recurrent, and video-based ANN models. Video models provided modest improvements in neural predictivity, particularly at later response stages, raising the question of what kind of dynamics they capture. To probe this, we applied a stress test: decoders trained on naturalistic videos were evaluated on "appearance-free" variants that preserve motion but remove shape and texture. IT population activity generalized across this manipulation, but all ANN classes failed. Thus, current video models better capture appearance-bound dynamics rather than the appearance-invariant temporal computations expressed in IT, underscoring the need for new objectives that encode biological temporal statistics and invariances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。