无需训练即可实现视频语义混合,保持对象一致性与时间连贯性。
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
- 基于无类别分割和对角去噪调度,定位并追踪视频中目标物体。
- 在语义变化下仍保持稳定,提升前后帧一致性,显著优于现有方法。
- 适合需要快速、可控视频编辑的创作者或研究者使用。
我们提出 MoCA-Video,一种无需训练的视频语义混合框架。该方法在冻结的视频扩散模型的潜在空间中运行,利用无类别分割结合对角去噪调度器,在帧间精准定位并追踪目标对象。为确保语义变化下的时间稳定性,引入基于动量的修正机制以近似训练数据分布之外的新混合分布,并采用轻量级伽马残差模块消除视觉伪影。通过结构相似性(SSIM)、感知图像质量(LPIPS)以及自定义度量 extmetricnameabbr(量化参考与输出间的语义对齐程度)进行评估。大量实验表明,本模型持续优于训练自由及有训练基线,在无需重训的前提下实现更优的语义混合效果与时间连贯性。结果证明,对扩散噪声轨迹的结构化操控可实现可控且高质量的语义变化视频编辑。
原文摘要 · Abstract (English)
We present MoCA-Video, a training-free framework for semantic mixing in videos. Operating in the latent space of a frozen video diffusion model, MoCA-Video utilizes class-agnostic segmentation with diagonal denoising scheduler to localize and track the target object across frames. To ensure temporal stability under semantic shifts, we introduce momentum-based correction to approximate novel hybrid distributions beyond trained data distribution, alongside a light gamma residual module that smooths out visual artifacts. We evaluate model's performance using SSIM, LPIPS, and a proposed metric, \metricnameabbr, which quantifies semantic alignment between reference and output. Extensive evaluation demonstrates that our model consistently outperforms both training-free and trained baselines, achieving superior semantic mixing and temporal coherence without retraining. Results establish that structured manipulation of diffusion noise trajectories enables controllable and high-quality video editing under semantic shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。