arXiv:2606.19676cs.CVcs.AI2026-06

一键同步编辑视频动作与位置,提升生成质量与可控性。

TeleMorpher: Toward Robust Simultaneous Motion-Location Editing

论文配图:TeleMorpher: Toward Robust Simultaneous Motion-Location Editing
图 1 · 摘自论文原文
  • 利用动作先验和目标动作引导,实现动作与位置的联合编辑
  • 在真实视频和太极数据集上均优于现有方法,定量与定性指标领先
  • 适合需要精准控制动作与位置变化的视频编辑场景

扩散模型在图像与视频生成编辑中取得显著进展。尽管近期研究拓展至动作编辑,但同时改变动作与位置——这一具有实际意义的任务——仍鲜有探索。为理解鲁棒的动作-位置编辑,我们首先分析了影响其质量的根本因素。基于此,提出TeleMorpher,据我们所知首个一次性框架,用于同步动作与位置编辑。该方法借助动作先验、由现成模型生成的目标动作中心视频作为编辑引导,以及真实动作信息,实现更可控、精确的编辑。流程包括:(1) 通过预训练分割与修复模型分离主角与背景;(2) 引入无训练姿态扭曲,以动作先验为指导编辑主角动作;(3) 将扭曲后视频直接注入基线动作编辑器进行推理,减少源与目标动作差异,同时保留源视频外观;(4) 为提升评估可靠性,提出两种基于LPIPS的新指标,分别衡量编辑前后背景一致性及动作编辑保真度(通过对比源与目标视频主角骨架差异)。在真实视频与TaiChi数据集上的实验表明,TeleMorpher在定量与定性评价(真人评估)中均表现优异,验证其有效性。

原文摘要 · Abstract (English)

Diffusion models have achieved remarkable success in image and video generation and editing. While recent studies have extended these efforts toward motion editing, simultaneously transforming both motion and location-despite its practical importance-remains largely unexplored. To better understand robust motion-location editing, we first analyze the fundamental factors that degrade its quality. Based on this analysis, we propose TeleMorpher, one of the first one-shot frameworks to the best of our knowledge, for simultaneous motion-location editing. Our approach leverages motion priors, a target motion-centric video generated from an off-the-shelf model as motion-editing guidance, and the ground truth motion to enable more controllable and precise motion-location editing. Via this, our framework works as follows: (1) we first disentangle the protagonist and the background via pre-trained segmentation and inpainting models. (2) Then, we introduce a training-free pose warping that edits the protagonist's motion with the motion prior as the guidance. (3) The result of warped motion video is directly injected into a baseline motion editor during inference, mitigating the difference between source and target motions while preserving the appearance of the source video. (4) To enhance the reliability of quantitative evaluations, we propose two new LPIPS-based metrics that measure the background consistency before and after the motion editing and the fidelity of motion editing performance via measuring the difference between the extracted protagonist's skeletons from source and target videos. Experiments with in-the-wild videos and the TaiChi dataset demonstrate that TeleMorpher achieves superior performance across both quantitative and qualitative measurements (real-human evaluation), underscoring its effectiveness.

视频编辑动作生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。