arXiv:2603.09084cs.CV2026-03

无需训练即可实现精准口型同步与音视频编辑

OmniEdit: A Training-free framework for Lip Synchronization and Audio-Visual Editing

  • 用目标序列替代编辑序列,实现无偏输出估计
  • 去除生成过程中的随机性,保证编辑轨迹平滑稳定
  • 适用于虚拟形象、影视制作等需高效音视频编辑场景

口型同步与音视频编辑是多模态学习中的基础挑战,广泛应用于影视制作、虚拟形象和远程通信等领域。尽管近期取得进展,现有方法大多依赖预训练模型的监督微调,带来巨大的计算开销和数据需求。本文提出OmniEdit,一种无需训练的框架,用于口型同步与音视频编辑。通过将FlowEdit中的编辑序列替换为目标序列,实现对期望输出的无偏估计;同时,消除生成过程中的随机因素,建立平滑稳定的编辑路径。大量实验验证了该框架的有效性与鲁棒性。代码已公开于https://github.com/l1346792580123/OmniEdit。

原文摘要 · Abstract (English)

Lip synchronization and audio-visual editing have emerged as fundamental challenges in multimodal learning, underpinning a wide range of applications, including film production, virtual avatars, and telepresence. Despite recent progress, most existing methods for lip synchronization and audio-visual editing depend on supervised fine-tuning of pre-trained models, leading to considerable computational overhead and data requirements. In this paper, we present OmniEdit, a training-free framework designed for both lip synchronization and audio-visual editing. Our approach reformulates the editing paradigm by substituting the edit sequence in FlowEdit with the target sequence, yielding an unbiased estimation of the desired output. Moreover, by removing stochastic elements from the generation process, we establish a smooth and stable editing trajectory. Extensive experimental results validate the effectiveness and robustness of the proposed framework. Code is available at https://github.com/l1346792580123/OmniEdit.

口型同步音视频编辑零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。