arXiv:2512.02015cs.CV2025-12被引 20

用3D点轨迹实现视频中相机与物体运动的精准联合编辑。

Generative Video Motion Editing with 3D Point Tracks

  • 通过3D点轨迹条件控制视频生成,实现精细运动编辑。
  • 相比2D轨迹,3D轨迹提供深度信息,解决遮挡和顺序问题。
  • 支持镜头联动、运动迁移、非刚性变形等多样编辑,适合影视创作。

相机与物体运动是视频叙事的核心。然而,在复杂物体运动下精确编辑这些运动仍具挑战性。现有图像到视频(I2V)方法缺乏完整场景上下文,导致视频编辑不一致;视频到视频(V2V)方法虽能实现视角变化或基础物体平移,但对细粒度运动控制有限。本文提出一种基于点轨迹的条件化V2V框架,可联合编辑相机与物体运动。通过源视频与配对的3D点轨迹(表示源与目标运动)作为条件,建立稀疏对应关系,将源视频的丰富上下文传递至新运动,同时保持时空一致性。关键在于,相较于2D轨迹,3D轨迹提供显式深度线索,使模型能够解析深度顺序并处理遮挡,实现精准运动编辑。模型在合成数据和真实数据上分两阶段训练,支持多种运动编辑,包括相机与物体协同操作、运动迁移、非刚性形变,为视频编辑开辟全新创作可能。

原文摘要 · Abstract (English)

Camera and object motions are central to a video's narrative. However, precisely editing these captured motions remains a significant challenge, especially under complex object movements. Current motion-controlled image-to-video (I2V) approaches often lack full-scene context for consistent video editing, while video-to-video (V2V) methods provide viewpoint changes or basic object translation, but offer limited control over fine-grained object motion. We present a track-conditioned V2V framework that enables joint editing of camera and object motion. We achieve this by conditioning a video generation model on a source video and paired 3D point tracks representing source and target motions. These 3D tracks establish sparse correspondences that transfer rich context from the source video to new motions while preserving spatiotemporal coherence. Crucially, compared to 2D tracks, 3D tracks provide explicit depth cues, allowing the model to resolve depth order and handle occlusions for precise motion editing. Trained in two stages on synthetic and real data, our model supports diverse motion edits, including joint camera/object manipulation, motion transfer, and non-rigid deformation, unlocking new creative potential in video editing.

视频编辑3D轨迹运动控制生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。