arXiv:2511.18277cs.CV2025-11

用稀疏关键点精准控制视频编辑中的运动一致性

Point-to-Point: Sparse Motion Guidance for Controllable Video Editing

  • 提出锚点标记(anchor tokens)表示视频动态,通过扩散模型先验提取关键运动轨迹
  • 仅用少量点轨迹实现跨场景编辑,运动保真度与编辑精度均优于现有方法
  • 适合需要精确控制人物动作的视频编辑任务,如影视剪辑、动画制作

准确保留编辑过程中主体的运动特性仍是视频编辑的核心挑战。现有方法在编辑效果与运动保真度之间常存在权衡,因其依赖的运动表征或过度拟合布局,或仅隐式定义。为此,我们重新审视基于点的运动表征。然而,在无人工标注的情况下,跨多样化视频场景识别有意义的点仍具挑战。为此,我们提出一种新运动表征——锚点标记(anchor tokens),利用视频扩散模型的丰富先验,捕捉最核心的运动模式。锚点标记通过少量信息丰富的点轨迹紧凑编码视频动态,并可灵活重定位以匹配新主体。这使我们的方法,Point-to-Point,在多样场景下具备强泛化能力。大量实验表明,锚点标记带来更可控且语义对齐的视频编辑,在编辑保真度与运动保真度上均表现更优。

原文摘要 · Abstract (English)

Accurately preserving motion while editing a subject remains a core challenge in video editing tasks. Existing methods often face a trade-off between edit and motion fidelity, as they rely on motion representations that are either overfitted to the layout or only implicitly defined. To overcome this limitation, we revisit point-based motion representation. However, identifying meaningful points remains challenging without human input, especially across diverse video scenarios. To address this, we propose a novel motion representation, anchor tokens, that capture the most essential motion patterns by leveraging the rich prior of a video diffusion model. Anchor tokens encode video dynamics compactly through a small number of informative point trajectories and can be flexibly relocated to align with new subjects. This allows our method, Point-to-Point, to generalize across diverse scenarios. Extensive experiments demonstrate that anchor tokens lead to more controllable and semantically aligned video edits, achieving superior performance in terms of edit and motion fidelity.

视频编辑运动控制扩散模型稀疏引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。