arXiv:2510.00050cs.MMcs.AI2025-10被引 12

实现视频中物体级音画同步编辑,支持增删改且保持结构完整。

Object-AVEdit: An Object-level Audio-Visual Editing Model

  • 基于逆向重构-再生框架,实现音画双模态物体级控制。
  • 在物体增删改任务上达到先进水平,音画语义对齐更精准。
  • 适合影视后期、音视频内容创作人员使用。

视频后期制作与影视领域对音画编辑有极高需求。尽管已有大量模型探索音视频编辑,但在物体级音画操作方面仍面临挑战:需在音频与视觉模态中实现物体的添加、替换和移除,同时保持源实例的结构信息。本文提出 extbf{Object-AVEdit},基于逆向重构-再生范式,实现物体级音画编辑。为提升物体可控性,构建了词到声音物体对齐的音频生成模型,弥合音频与视频生成模型间的物体控制鸿沟。同时,提出一种全局优化的逆向-再生编辑算法,确保逆向过程的信息保留与再生效果更优。大量实验表明,该模型在物体级音视频编辑任务中表现领先,实现了精细的音画语义对齐。此外,所开发的音频生成模型也达到先进性能。更多结果见项目页:https://gewu-lab.github.io/Object_AVEdit-website/。

原文摘要 · Abstract (English)

There is a high demand for audio-visual editing in video post-production and the film making field. While numerous models have explored audio and video editing, they struggle with object-level audio-visual operations. Specifically, object-level audio-visual editing requires the ability to perform object addition, replacement, and removal across both audio and visual modalities, while preserving the structural information of the source instances during the editing process. In this paper, we present \textbf{Object-AVEdit}, achieving the object-level audio-visual editing based on the inversion-regeneration paradigm. To achieve the object-level controllability during editing, we develop a word-to-sounding-object well-aligned audio generation model, bridging the gap in object-controllability between audio and current video generation models. Meanwhile, to achieve the better structural information preservation and object-level editing effect, we propose an inversion-regeneration holistically-optimized editing algorithm, ensuring both information retention during the inversion and better regeneration effect. Extensive experiments demonstrate that our editing model achieved advanced results in both audio-video object-level editing tasks with fine audio-visual semantic alignment. In addition, our developed audio generation model also achieved advanced performance. More results on our project page: https://gewu-lab.github.io/Object_AVEdit-website/.

音画编辑物体级控制多模态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。