arXiv:2608.23790cs.CVq-bio.NC2026-08

人类与猴脑如何在外观变化中保持视觉稳定?

Primate vision reveals a missing principle for robust dynamic AI

  • 通过对比人脑与猴脑视觉皮层,发现运动信息逐步融入物体表征
  • 预测性模型在外观改变时仍能保持识别准确率,最接近大脑反应
  • 揭示了运动信息动态整合是鲁棒视觉的核心机制,适合做AI模型改进

智能视觉系统如何在外观变化时,将物体外观与运动信息结合并保持鲁棒性?我们通过比较人类感知和猕猴下颞叶皮层(macaque IT)的神经活动,以及基于图像和视频的神经网络模型(涵盖识别、分割、光流处理和预测世界建模),发现时间整合提升了物体表征。然而,大多数视频识别模型在外观被破坏但运动结构保留时泛化能力差,而人类和猕猴IT仍保持鲁棒性。值得注意的是,预测性世界模型在跨外观泛化上表现最佳,且与IT神经活动的对应度最高,在神经保真度上优于其他视频建模方法。但无一模型能再现皮层从早期以外观为主到后期出现外观不变运动编码的动态转换。这些结果表明,运动信息逐步整合进物体表征是鲁棒动态视觉的关键原则,并提示预测学习可能是实现这一计算的人工系统可行路径。

原文摘要 · Abstract (English)

How does an intelligent visual system combine what objects look like with how they move while remaining robust as appearance changes? We addressed this question by comparing human perception and neural activity in macaque inferior temporal cortex with representations from image- and video-based neural networks spanning recognition, segmentation, optic-flow processing and predictive world modeling. Temporal integration improved object representations, but most video recognition models generalized poorly when appearance was disrupted while motion structure was preserved. Humans and macaque IT remained robust. Notably, predictive world models combined strong cross-appearance generalization with the closest correspondence to IT, outperforming other video-modeling approaches in neural fidelity. Yet no model reproduced the cortical transformation from early appearance-dominated responses toward later appearance-invariant motion coding. These results identify progressive integration of motion into object representations as a principle of robust dynamic vision and implicate predictive learning as a promising route toward realizing this computation in artificial systems.

动态视觉预测模型神经机制鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。