arXiv:2501.00645cs.CVcs.LG2025-01

用声音当画笔,直接操控视觉场景的生成与编辑。

SoundBrush: Sound as a Brush for Visual Scene Editing

  • 将音频特征映射到文本空间,实现声音驱动的图像编辑。
  • 可精准调整场景整体布局或插入发声物体,保持原始内容不变。
  • 支持3D场景编辑,适合影视、游戏等需要声画联动的领域。

我们提出 SoundBrush,一种以声音为画笔编辑和操控视觉场景的模型。通过扩展潜在扩散模型(LDM)的生成能力,引入音频信息实现视觉场景编辑。受现有图像编辑方法启发,我们将该任务建模为监督学习问题,并利用多种现成模型构建了音视频配对的数据集用于训练。该丰富数据集使 SoundBrush 能够学习将音频特征映射至 LDM 的文本空间,从而实现由真实世界声音引导的视觉场景编辑。与现有方法不同,SoundBrush 可准确操控整体场景甚至插入发声物体,以最佳匹配音频输入,同时保留原有内容。此外,结合新视角合成技术,本框架可拓展至3D场景编辑,实现声音驱动的3D场景操作。演示见 https://soundbrush.github.io/。

原文摘要 · Abstract (English)

We propose SoundBrush, a model that uses sound as a brush to edit and manipulate visual scenes. We extend the generative capabilities of the Latent Diffusion Model (LDM) to incorporate audio information for editing visual scenes. Inspired by existing image-editing works, we frame this task as a supervised learning problem and leverage various off-the-shelf models to construct a sound-paired visual scene dataset for training. This richly generated dataset enables SoundBrush to learn to map audio features into the textual space of the LDM, allowing for visual scene editing guided by diverse in-the-wild sound. Unlike existing methods, SoundBrush can accurately manipulate the overall scenery or even insert sounding objects to best match the audio inputs while preserving the original content. Furthermore, by integrating with novel view synthesis techniques, our framework can be extended to edit 3D scenes, facilitating sound-driven 3D scene manipulation. Demos are available at https://soundbrush.github.io/.

声音编辑图像生成3D编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。