用神经科学模型实现随机点运动下的类人零样本分割
Object segmentation from common fate: Motion energy processing enables human-like zero-shot generalization to random dot stimuli
- 基于脑区MT的运动能量模型处理运动能量信息
- 该模型在随机点视频中分割表现远超40种深度光流模型
- 首次实现计算机视觉对随机点刺激的类人零样本泛化
人类能根据“共同命运”原理,从运动中识别并分割物体。以往研究发现,人类感知可零样本推广至未见过的纹理或随机点。本文评估多种光流模型与受神经科学启发的运动能量模型在随机点刺激上的零样本图底分割能力。采用1998年Simoncelli和Heeger提出的、基于皮层区域MT神经记录拟合的运动能量模型。结果表明,40种在不同数据集上训练的深度光流模型均难以估计随机点视频中的运动模式,分割性能差;而运动能量模型显著优于所有光流模型。通过形状识别任务的心理物理学实验对比人类表现,所有先进光流模型均未达人类水平,唯独运动能量模型与人类能力相当。该模型成功弥补了当前计算机视觉在随机点刺激下类人零样本泛化能力的缺失,建立了格式塔心理学与大脑运动皮层处理之间的有力关联。代码、模型与数据集见https://github.com/mtangemann/motion_energy_segmentation。
原文摘要 · Abstract (English)
Humans excel at detecting and segmenting moving objects according to the Gestalt principle of "common fate". Remarkably, previous works have shown that human perception generalizes this principle in a zero-shot fashion to unseen textures or random dots. In this work, we seek to better understand the computational basis for this capability by evaluating a broad range of optical flow models and a neuroscience inspired motion energy model for zero-shot figure-ground segmentation of random dot stimuli. Specifically, we use the extensively validated motion energy model proposed by Simoncelli and Heeger in 1998 which is fitted to neural recordings in cortex area MT. We find that a cross section of 40 deep optical flow models trained on different datasets struggle to estimate motion patterns in random dot videos, resulting in poor figure-ground segmentation performance. Conversely, the neuroscience-inspired model significantly outperforms all optical flow models on this task. For a direct comparison to human perception, we conduct a psychophysical study using a shape identification task as a proxy to measure human segmentation performance. All state-of-the-art optical flow models fall short of human performance, but only the motion energy model matches human capability. This neuroscience-inspired model successfully addresses the lack of human-like zero-shot generalization to random dot stimuli in current computer vision models, and thus establishes a compelling link between the Gestalt psychology of human object perception and cortical motion processing in the brain. Code, models and datasets are available at https://github.com/mtangemann/motion_energy_segmentation
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。