用视频运动信息训练模型识别物体,无需标注也能学出高质量图像表示。
Object Concepts Emerge from Motion

- 利用光流和聚类生成伪实例掩码,指导单图编码器学习物体级表征
- 在42100万帧上训练,模型在几何与实例敏感任务上表现超越监督/自监督基线
- 完全无标注、无相机参数依赖,适合大规模视觉预训练场景
物体中心的视觉表征对物理世界感知至关重要,但现有视觉预训练方法常忽略个体实例的身份与连贯性。本文提出一种受生物启发的框架,从原始视频中学习单张图像的物体中心表征。方法以运动边界为物体分组信号:通过现成光流与聚类生成伪实例掩码,监督单图编码器进行像素级成对度量学习。该框架无需人工标注或相机标定。我们首先从7,163小时驾驶与网络视频中获取1.95亿个伪标注帧,再通过运动验证的自训练扩展至4.21亿帧。训练了最大至Swin-H的编码器,并将学到的表征蒸馏到一系列Swin主干网络。在单目深度估计、3D目标检测、3D占据预测及端到端规划任务中,模型性能达到或优于监督与自监督基线,尤其在几何与实例敏感任务上转移能力突出。结果表明,运动衍生的监督可使静态图像编码器学会表征视觉实例,为可扩展视觉预训练提供了互补路径。
原文摘要 · Abstract (English)
Object-centric visual representations are important for physical-world perception, but existing visual pretraining methods often capture semantic categories without preserving the identity and coherence of individual instances. We present a biologically inspired framework that learns object-centric representations for single images from raw videos. Our approach uses motion boundaries as a source of object-level grouping: off-the-shelf optical flow and clustering produce pseudo-instance masks, which supervise a single-image encoder with pixel-level pairwise metric learning. The framework requires neither human annotations nor camera calibration. We first obtain 195 million pseudo-labeled frames from 7,163 hours of driving and web videos, then expand the supervision to 421 million frames with Motion-Verified Self-Training, which combines model proposals with motion evidence. We train encoders up to Swin-H and distill the learned representations into a family of Swin backbones. Across monocular depth estimation, 3D object detection, 3D occupancy prediction, and end-to-end planning, the resulting models achieve competitive or superior performance relative to supervised and self-supervised pretraining baselines, with particularly strong transfer on geometry- and instance-sensitive tasks. These results show that motion-derived supervision can teach static image encoders to represent visual instances, providing a complementary direction for scalable visual pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。