arXiv:2504.18283cs.CVcs.AI2025-04

用混音音频生成图像并分离出不同声源对应的画面。

Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator

  • 用音视频分离器从混合音频中解析出多类声音,分别生成对应图像。
  • 在VGGSound数据集上,生成图像的类表征分数提升7%,召回率提高4%。
  • 首次提出多类音频生成与分离任务,适合音视频生成与内容理解研究者。

近年来,音视频生成模型在从单类音频生成图像方面取得显著进展。然而,现有方法仅能处理单一类别音频,无法生成来自混合音频的图像。为此,我们提出音视频生成与分离模型(AV-GAS),用于从声景(包含多个类别的混合音频)生成图像。贡献有三:第一,提出新挑战——给定多类音频输入生成图像,并通过音视频分离器实现;第二,引入新任务:为混合音频中每类声音生成独立图像;第三,提出新评估指标:类表征得分(CRS)和改进的R@K。模型在VGGSound数据集上训练与评估,结果表明该方法优于现有最先进水平,在生成合理图像时达到7%更高的CRS和4%更高的R@2*。

原文摘要 · Abstract (English)

Recent audio-visual generative models have made substantial progress in generating images from audio. However, existing approaches focus on generating images from single-class audio and fail to generate images from mixed audio. To address this, we propose an Audio-Visual Generation and Separation model (AV-GAS) for generating images from soundscapes (mixed audio containing multiple classes). Our contribution is threefold: First, we propose a new challenge in the audio-visual generation task, which is to generate an image given a multi-class audio input, and we propose a method that solves this task using an audio-visual separator. Second, we introduce a new audio-visual separation task, which involves generating separate images for each class present in a mixed audio input. Lastly, we propose new evaluation metrics for the audio-visual generation task: Class Representation Score (CRS) and a modified R@K. Our model is trained and evaluated on the VGGSound dataset. We show that our method outperforms the state-of-the-art, achieving 7% higher CRS and 4% higher R@2* in generating plausible images with mixed audio.

音视频生成声景分析图像分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。