arXiv:2508.20379cs.CV2025-08中稿 · BMVC 2025

用音频+文字联合指导图像编辑,无需训练即可处理复杂场景。

Audio-Guided Visual Editing with Complex Multi-Modal Prompts

  • 通过多模态编码器融合文本与音频,实现零样本跨模态对齐。
  • 采用独立噪声分支与自适应补丁选择,支持多源多模态提示。
  • 在复杂编辑任务中表现优于纯文本方法,适合真实世界应用。

基于扩散模型的图像编辑虽有进展,但在需多维度描述的复杂场景中常因仅依赖文本提示而效果不佳,亟需非文本提示补充。本文提出一种无需额外训练的音频引导图像编辑框架,利用具备强零样本能力的预训练多模态编码器,将多样音频融入编辑过程,缓解音频编码空间与扩散模型提示编码空间间的差异。同时,提出新的分离噪声分支与自适应补丁选择机制,有效处理多模态、多提示的复杂编辑任务。在多种编辑任务上的实验表明,该框架能充分整合音频提供的丰富信息,在文本无法胜任的复杂场景中表现优异。

原文摘要 · Abstract (English)

Visual editing with diffusion models has made significant progress but often struggles with complex scenarios that textual guidance alone could not adequately describe, highlighting the need for additional non-text editing prompts. In this work, we introduce a novel audio-guided visual editing framework that can handle complex editing tasks with multiple text and audio prompts without requiring additional training. Existing audio-guided visual editing methods often necessitate training on specific datasets to align audio with text, limiting their generalization to real-world situations. We leverage a pre-trained multi-modal encoder with strong zero-shot capabilities and integrate diverse audio into visual editing tasks, by alleviating the discrepancy between the audio encoder space and the diffusion model's prompt encoder space. Additionally, we propose a novel approach to handle complex scenarios with multiple and multi-modal editing prompts through our separate noise branching and adaptive patch selection. Our comprehensive experiments on diverse editing tasks demonstrate that our framework excels in handling complicated editing scenarios by incorporating rich information from audio, where text-only approaches fail.

图像编辑多模态音频引导扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。