用自然视频同时学识别与运动,突破单任务自监督学习瓶颈。
Midway Network: Learning Representations for Recognition and Motion from Latent Dynamics
- 通过中间层自上而下路径推断帧间运动隐变量
- 在两个大规模数据集预训练后,分割与光流任务表现领先
- 首次实现对高阶对应关系的动态捕捉分析
物体识别与运动理解是互补的感知核心。尽管自监督学习在无标签数据中展现潜力,但现有方法多聚焦于识别或运动任一任务。本工作提出Midway Network,首个仅基于自然视频、同时学习强视觉表征用于识别与运动理解的自监督架构。其通过中间层自上而下路径推断帧间运动隐变量,并结合密集前向预测目标与分层结构,应对自然视频中复杂的多对象场景。在两个大规模自然视频数据集上预训练后,模型在语义分割和光流任务上均优于以往自监督方法。此外,我们提出基于前向特征扰动的新分析方法,验证了模型所学动态可捕捉高层对应关系。
原文摘要 · Abstract (English)
Object recognition and motion understanding are key components of perception that complement each other. While self-supervised learning methods have shown promise in their ability to learn from unlabeled data, they have primarily focused on obtaining rich representations for either recognition or motion rather than both in tandem. On the other hand, latent dynamics modeling has been used in decision making to learn latent representations of observations and their transformations over time for control and planning tasks. In this work, we present Midway Network, a new self-supervised learning architecture that is the first to learn strong visual representations for both object recognition and motion understanding solely from natural videos, by extending latent dynamics modeling to this domain. Midway Network leverages a midway top-down path to infer motion latents between video frames, as well as a dense forward prediction objective and hierarchical structure to tackle the complex, multi-object scenes of natural videos. We demonstrate that after pretraining on two large-scale natural video datasets, Midway Network achieves strong performance on both semantic segmentation and optical flow tasks relative to prior self-supervised learning methods. We also show that Midway Network's learned dynamics can capture high-level correspondence via a novel analysis method based on forward feature perturbation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。