测试视觉模型能否像人一样感知静态图像中的错觉运动。
Do vision models perceive illusory motion in static images like humans?

- 用眼动模拟测试不同光流模型对旋转蛇错觉的响应。
- 只有受人类视觉启发的双通道模型在眼动模拟下产生旋转流动。
- 递归注意力机制和亮度/颜色特征是关键,适合做类人视觉系统。
理解人类运动感知对构建以人为中心的计算机视觉系统至关重要。尽管深度神经网络(DNN)在光流估计上表现强劲,但其鲁棒性仍逊于人类,且依赖根本不同的计算策略。视觉运动错觉为探测这些机制提供了有力工具,揭示了人机视觉的异同。虽然近期基于DNN的运动模型能复现动态错觉(如反向菲亚),但尚不清楚它们是否能在静态图像中感知错觉运动,例如旋转蛇错觉。我们评估了几种代表性光流模型在旋转蛇错觉上的表现,发现大多数模型无法生成与人类感知一致的流动场。在模拟扫视眼动的条件下,仅人类启发的双通道模型表现出预期的旋转运动,且在扫视模拟期间对应关系最接近。消融分析进一步表明,亮度基信号和高阶颜色-特征基信号均对此行为有贡献,且递归注意力机制对于整合局部线索至关重要。结果凸显当前光流模型与人类视觉运动处理之间的显著差距,并为未来更贴近人类感知的运动估计系统和以人为本的AI提供洞见。
原文摘要 · Abstract (English)
Understanding human motion processing is essential for building reliable, human-centered computer vision systems. Although deep neural networks (DNNs) achieve strong performance in optical flow estimation, they remain less robust than humans and rely on fundamentally different computational strategies. Visual motion illusions provide a powerful probe into these mechanisms, revealing how human and machine vision align or diverge. While recent DNN-based motion models can reproduce dynamic illusions such as reverse-phi, it remains unclear whether they can perceive illusory motion in static images, exemplified by the Rotating Snakes illusion. We evaluate several representative optical flow models on Rotating Snakes and show that most fail to generate flow fields consistent with human perception. Under simulated conditions mimicking saccadic eye movements, only the human-inspired Dual-Channel model exhibits the expected rotational motion, with the closest correspondence emerging during the saccade simulation. Ablation analyses further reveal that both luminance-based and higher-order color--feature--based motion signals contribute to this behavior and that a recurrent attention mechanism is critical for integrating local cues. Our results highlight a substantial gap between current optical-flow models and human visual motion processing, and offer insights for developing future motion-estimation systems with improved correspondence to human perception and human-centric AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。