arXiv:2512.12875cs.CVcs.MM2025-12被引 2

实现音视频中目标对象的同步移除,保持内容连贯性。

Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal

  • 基于流匹配的端到端模型,同步编辑音视频
  • 在保留其余内容的前提下精准移除指定对象
  • 适合需要音视频协同编辑的创作者使用

音视频联合编辑对精确可控的内容创作至关重要。由于编辑前后配对音视频数据的缺失以及模态间的异质性,这一任务面临挑战。为此,我们构建了SAVEBench数据集,包含文本和掩码条件下的配对音视频数据,支持基于对象的源到目标学习。基于此,我们训练了Schrodinger Audio-Visual Editor(SAVE),一个端到端的流匹配模型,在处理过程中保持音视频同步。SAVE引入了薛定谔桥机制,直接学习从源到目标音视频混合物的传输。评估表明,该模型能够有效移除目标对象,同时保持更强的时间同步性和音视频语义一致性,优于音视频编辑器的简单组合。

原文摘要 · Abstract (English)

Joint editing of audio and visual content is crucial for precise and controllable content creation. This new task poses challenges due to the limitations of paired audio-visual data before and after targeted edits, and the heterogeneity across modalities. To address the data and modeling challenges in joint audio-visual editing, we introduce SAVEBench, a paired audiovisual dataset with text and mask conditions to enable object-grounded source-to-target learning. With SAVEBench, we train the Schrodinger Audio-Visual Editor (SAVE), an end-to-end flow-matching model that edits audio and video in parallel while keeping them aligned throughout processing. SAVE incorporates a Schrodinger Bridge that learns a direct transport from source to target audiovisual mixtures. Our evaluation demonstrates that the proposed SAVE model is able to remove the target objects in audio and visual content while preserving the remaining content, with stronger temporal synchronization and audiovisual semantic correspondence compared with pairwise combinations of an audio editor and a video editor.

音视频编辑跨模态生成同步生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。