arXiv:2510.12069cs.CV2025-10ICCV

用姿态和位置先验实现动作保持的视频编辑,支持灵活结构与语义替换。

VIDMP3: Video Editing by Representing Motion with Pose and Position Priors

  • 基于源视频学习姿态与位置先验,构建通用运动表征。
  • 生成视频保持原始动作,且在结构语义上可自由调整。
  • 无需人工干预,解决时序不一致与主体身份漂移问题。

动作保持的视频编辑对创作者至关重要,尤其在需要灵活调整对象结构与语义的场景中。尽管潜力巨大,该领域仍研究不足。现有基于扩散模型的编辑方法在结构保持任务中表现优异,依赖密集引导信号确保内容完整性。部分近期方法尝试解决结构可变编辑问题,但常出现时序不一致、主体身份漂移,或需人工干预。为此,我们提出VidMP3,通过姿态与位置先验从源视频中学习通用运动表征,使生成视频在保持原始动作的同时,支持结构与语义的灵活性。定性与定量评估均证明该方法优于现有方法。代码将公开于https://github.com/sandeep-sm/VidMP3。

原文摘要 · Abstract (English)

Motion-preserved video editing is crucial for creators, particularly in scenarios that demand flexibility in both the structure and semantics of swapped objects. Despite its potential, this area remains underexplored. Existing diffusion-based editing methods excel in structure-preserving tasks, using dense guidance signals to ensure content integrity. While some recent methods attempt to address structure-variable editing, they often suffer from issues such as temporal inconsistency, subject identity drift, and the need for human intervention. To address these challenges, we introduce VidMP3, a novel approach that leverages pose and position priors to learn a generalized motion representation from source videos. Our method enables the generation of new videos that maintain the original motion while allowing for structural and semantic flexibility. Both qualitative and quantitative evaluations demonstrate the superiority of our approach over existing methods. The code will be made publicly available at https://github.com/sandeep-sm/VidMP3.

视频编辑运动表征扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。