arXiv:2410.07463cs.CV2024-10被引 23

用一句话描述声音和画面,实现精准音画编辑。

Language-Guided Joint Audio-Visual Editing via One-Shot Adaptation

  • 仅需一个样本即可适配音画生成模型,实现快速定制。
  • 通过语言引导编辑,保持音画内容一致性,避免视觉忽略。
  • 适合音视频创作、虚拟场景构建等需要精准控制的场景。

本文提出一种新型任务:语言引导的联合音画编辑。给定一个声音与图像对应的声音事件,该任务旨在根据语言指令生成新的音画内容。例如,在不改变物体外观的前提下更换背景环境,或为视觉内容添加符合语境的新声音。为此,我们提出基于扩散模型的联合音画编辑框架,并引入两项关键创新:首先,提出单样本适配方法,仅需一个音画样本即可同时将音频与视觉扩散模型迁移至目标域,微调后可一致生成该样本;其次,提出跨模态语义增强策略,缓解语言引导编辑中视觉分支对编辑需求的“灾难性忽视”问题,提升语言与视觉间的语义一致性。大量实验验证了方法在语言引导音画编辑中的有效性,显著优于多个基线方法。更多信息请访问项目页:https://liangsusan-git.github.io/project/avedit/

原文摘要 · Abstract (English)

In this paper, we introduce a novel task called language-guided joint audio-visual editing. Given an audio and image pair of a sounding event, this task aims at generating new audio-visual content by editing the given sounding event conditioned on the language guidance. For instance, we can alter the background environment of a sounding object while keeping its appearance unchanged, or we can add new sounds contextualized to the visual content. To address this task, we propose a new diffusion-based framework for joint audio-visual editing and introduce two key ideas. Firstly, we propose a one-shot adaptation approach to tailor generative diffusion models for audio-visual content editing. With as few as one audio-visual sample, we jointly transfer the audio and vision diffusion models to the target domain. After fine-tuning, our model enables consistent generation of this audio-visual sample. Secondly, we introduce a cross-modal semantic enhancement approach. We observe that when using language as content editing guidance, the vision branch may overlook editing requirements. This phenomenon, termed catastrophic neglect, hampers audio-visual alignment during content editing. We therefore enhance semantic consistency between language and vision to mitigate this issue. Extensive experiments validate the effectiveness of our method in language-based audio-visual editing and highlight its superiority over several baseline approaches. We recommend that readers visit our project page for more details: https://liangsusan-git.github.io/project/avedit/.

音画编辑扩散模型语言引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。