arXiv:2607.21576cs.CV2026-07

从视频中分离物体运动与相机运动,提升动作表征的鲁棒性。

Self-Supervised Learning of Structured Dynamics from Videos

论文配图:Self-Supervised Learning of Structured Dynamics from Videos
图 1 · 摘自论文原文
  • 通过未来特征预测显式分离主导动态与残差动态
  • 在真实与合成视频上优于基线模型,接近强监督表现
  • 适用于需要解耦运动因素的视觉分析任务

理解视频中的运动是视觉学习的基础挑战,因帧间变化同时包含相机运动和物体运动。这种分解在表征学习中仍被忽视,主要因为二者在自然视频中紧密耦合且难以独立标注。但分离它们对学习区分有意义物体动态与相机干扰的鲁棒表征至关重要。本文研究能否从预训练图像Transformer的冻结特征中恢复结构化运动表征。提出结构化动态模型(SDM),通过未来特征预测而非单一纠缠隐状态或无结构的空间密集转移标记,显式分离主导时序变化与残差动态。训练结合真实视频的自监督学习与合成Kubric数据上的弱监督场景动态。在新提出的ProbeMotion评估套件上测试,该套件涵盖合成与真实视频中的相机运动、物体运动及混合动态。SDM优于使用全局CLS或平均池化特征的基线模型,在多个探针任务上媲美强监督表示如VGGT,尽管监督强度显著更弱。结果表明,预训练图像模型可直接重用于结构化视频动态表征,为学习和分析潜在视频动态提供有效归纳偏置。

原文摘要 · Abstract (English)

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of dynamics: camera motion and object motion. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object dynamics from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Dynamics Model (SDM), which explicitly separates the dominant source of temporal change from residual dynamics through future-feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene dynamics on synthetic Kubric data. We evaluate SDM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera motion, object motion, and combined dynamics. SDM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured video-dynamics representations, providing a useful inductive bias for learning and analyzing latent video dynamics.

视频生成自监督学习动态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。