用户可选图像物体,模型精准生成对应声音。
Sounding that Object: Interactive Object-Aware Image to Audio Generation
- 基于多模态注意力,让声音与图像物体对齐。
- 测试时通过分割实现物体级音效生成,效果优于基线。
- 适合交互式音视频生成、虚拟场景构建等应用。
在包含多个物体和声源的复杂视听场景中,生成准确的声音极具挑战性。本文提出一种交互式物体感知音频生成模型,将声音生成锚定在用户选定的图像物体上。方法将物体中心学习融入条件潜空间扩散模型,通过多模态注意力学习图像区域与其对应声音的关联。测试时,利用图像分割技术实现用户对物体级别的声音交互生成。理论证明,其注意力机制在功能上近似于测试时的分割掩码,确保生成音频与所选物体对齐。定量与定性评估均表明,该模型在物体与声音匹配度上优于基线,显著提升音画一致性。
原文摘要 · Abstract (English)
Generating accurate sounds for complex audio-visual scenes is challenging, especially in the presence of multiple objects and sound sources. In this paper, we propose an {\em interactive object-aware audio generation} model that grounds sound generation in user-selected visual objects within images. Our method integrates object-centric learning into a conditional latent diffusion model, which learns to associate image regions with their corresponding sounds through multi-modal attention. At test time, our model employs image segmentation to allow users to interactively generate sounds at the {\em object} level. We theoretically validate that our attention mechanism functionally approximates test-time segmentation masks, ensuring the generated audio aligns with selected objects. Quantitative and qualitative evaluations show that our model outperforms baselines, achieving better alignment between objects and their associated sounds. Project page: https://tinglok.netlify.app/files/avobject/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。