无需标注数据,通过反事实优化实现视频运动估计新方法。
Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals

- 基于反事实探针优化,从预训练模型中提取运动信息
- 在真实视频上达到当前最优运动估计性能
- 适用于无监督视频分析与可控视频生成场景
视频运动估计是计算机视觉中的关键问题,广泛应用于可控视频生成和机器人等领域。现有方法多依赖合成数据或特定场景的启发式调参,限制了其在真实世界中的表现。尽管大规模自监督视频学习取得进展,但将其用于运动估计仍不充分。本文提出 Opt-CWM,一种基于预训练帧预测模型的自监督运动与遮挡估计方法。该方法通过优化反事实探针,从基础视频模型中提取运动信息,无需固定启发式规则,且可在无限制视频输入下训练。在真实视频上实现了当前最佳的运动估计性能,且完全无需标签数据。
原文摘要 · Abstract (English)
Estimating motion in videos is an essential computer vision problem with many downstream applications, including controllable video generation and robotics. Current solutions are primarily trained using synthetic data or require tuning of situation-specific heuristics, which inherently limits these models' capabilities in real-world contexts. Despite recent developments in large-scale self-supervised learning from videos, leveraging such representations for motion estimation remains relatively underexplored. In this work, we develop Opt-CWM, a self-supervised technique for flow and occlusion estimation from a pre-trained next-frame prediction model. Opt-CWM works by learning to optimize counterfactual probes that extract motion information from a base video model, avoiding the need for fixed heuristics while training on unrestricted video inputs. We achieve state-of-the-art performance for motion estimation on real-world videos while requiring no labeled data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。