用生成模型模拟人眼感知物体运动,统一处理不同场景下的物体识别。
GenMatter: Perceiving Physical Objects with Generative Matter Models

- 分层建模运动与外观特征,将像素聚为粒子再聚为可动物体
- 在随机点、纹理和自然视频中均实现与人眼一致的物体感知
- 适合研究视觉机制或需鲁棒运动理解的场景理解任务
人类视觉能稳健地检测并分割由独立运动物质构成的实体,无论面对稀疏运动点、纹理表面还是自然场景。现有计算机视觉系统缺乏跨多种场景的统一方法。受人类感知原理启发,我们提出一种生成模型,将低层运动线索与高层外观特征分层聚合为粒子(代表局部物质的小高斯),再将粒子聚为簇以捕捉协同且独立运动的物理实体。我们开发了基于并行化块吉布斯采样的硬件加速推理算法,实现稳定粒子运动与分组恢复。该模型可处理随机点、风格化纹理或自然RGB视频,适用于生物视觉成功而传统方法失败的多种场景。我们在三个领域验证:在2D随机点动画中,模型捕捉到人类对模糊条件下的分级不确定性;在类格式塔的伪装旋转物体数据集上,模型从运动恢复正确3D结构并实现准确2D分割;在自然视频中,模型追踪变形物体的三维运动物质,支持鲁棒的物体级场景理解。本工作建立了一个基于人类视觉原则的通用运动感知框架。
原文摘要 · Abstract (English)
Human visual perception offers valuable insights for understanding computational principles of motion-based scene interpretation. Humans robustly detect and segment moving entities that constitute independently moveable chunks of matter, whether observing sparse moving dots, textured surfaces, or naturalistic scenes. In contrast, existing computer vision systems lack a unified approach that works across these diverse settings. Inspired by principles of human perception, we propose a generative model that hierarchically groups low-level motion cues and high-level appearance features into particles (small Gaussians representing local matter), and groups particles into clusters capturing coherently and independently moveable physical entities. We develop a hardware-accelerated inference algorithm based on parallelized block Gibbs sampling to recover stable particle motion and groupings. Our model operates on different kinds of inputs (random dots, stylized textures, or naturalistic RGB video), enabling it to work across settings where biological vision succeeds but existing computer vision approaches do not. We validate this unified framework across three domains: on 2D random dot kinematograms, our approach captures human object perception including graded uncertainty across ambiguous conditions; on a Gestalt-inspired dataset of camouflaged rotating objects, our approach recovers correct 3D structure from motion and thereby accurate 2D object segmentation; and on naturalistic RGB videos, our model tracks the moving 3D matter that makes up deforming objects, enabling robust object-level scene understanding. This work thus establishes a general framework for motion-based perception grounded in principles of human vision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。