arXiv:2512.07209cs.MMcs.LG2025-12被引 2

视频修改后自动匹配同步音频,让视听更连贯。

Coherent Audio-Visual Editing via Conditional Audio Generation Following Video Edits

  • 用视频和文本控制生成新音频,保持音画一致。
  • 在复杂修改时智能保留原音频结构,效果优于现有方法。
  • 适合影视后期、虚拟内容生成等需要音画协同的场景。

我们提出一种新的音画联合编辑流程,提升修改后视频与伴随音频的一致性。首先使用先进的视频编辑技术生成目标视频,再进行音频编辑以匹配视觉变化。为此,我们设计了一种新型视频到音频生成模型,该模型以源音频、目标视频和文本提示为条件输入。通过扩展模型架构引入条件音频输入,并提出一种数据增强策略以提升训练效率。此外,模型能根据编辑复杂度动态调整源音频的影响程度,在可能情况下保留原始音频结构。实验表明,该方法在维持音画对齐和内容完整性方面优于现有技术。

原文摘要 · Abstract (English)

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then performs audio editing to align with the visual changes. To achieve this, we present a new video-to-audio generation model that conditions on the source audio, target video, and a text prompt. We extend the model architecture to incorporate conditional audio input and propose a data augmentation strategy that improves training efficiency. Furthermore, our model dynamically adjusts the influence of the source audio based on the complexity of the edits, preserving the original audio structure where possible. Experimental results demonstrate that our method outperforms existing approaches in maintaining audio-visual alignment and content integrity.

音画同步视频编辑条件生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。