arXiv:2503.10287cs.SDcs.CV2025-03AAAI被引 3

首次实现多源音频生成图像,提升视觉内容完整性。

MACS: Multi-source Audio-to-image Generation with Contextual Significance and Semantic Alignment

  • 分两阶段处理:先分离多源音频,再映射生成图像。
  • 在21项指标中17项超越现有方法,生成图像质量更优。
  • 适合跨模态生成、音频理解与多源信号处理研究者。

得益于深度生成模型的突破,音频到图像生成已成为关键的跨模态任务,将复杂听觉信号转化为丰富的视觉表征。然而,以往工作仅关注单源音频输入,忽略了自然听觉场景中的多源特性,限制了生成内容的全面性。为此,我们提出MACS方法,首次显式分离多源音频以捕捉丰富音频成分。MACS为两阶段方法:第一阶段通过弱监督方式分离多源音频,利用预训练的CLAP模型将音频与文本标签映射至同一语义空间,并引入排序损失以考虑分离音频的上下文重要性;第二阶段通过可训练适配器和MLP层,将分离后的音频信号映射为图像生成条件。我们对LLP数据集进行预处理,构建首个完整的多源音频到图像生成基准。实验在多源、混合源及单源任务上展开,结果显示MACS在21项评估指标中有17项优于当前最优方法,且生成图像质量显著提升。

原文摘要 · Abstract (English)

Propelled by the breakthrough in deep generative models, audio-to-image generation has emerged as a pivotal cross-modal task that converts complex auditory signals into rich visual representations. However, previous works only focus on single-source audio inputs for image generation, ignoring the multi-source characteristic in natural auditory scenes, thus limiting the performance in generating comprehensive visual content. To bridge this gap, we propose a method called MACS to conduct multi-source audio-to-image generation. To our best knowledge, this is the first work that explicitly separates multi-source audio to capture the rich audio components before image generation. MACS is a two-stage method. In the first stage, multi-source audio inputs are separated by a weakly supervised method, where the audio and text labels are semantically aligned by casting into a common space using the large pre-trained CLAP model. We introduce a ranking loss to consider the contextual significance of the separated audio signals. In the second stage, effective image generation is achieved by mapping the separated audio signals to the generation condition using only a trainable adapter and a MLP layer. We preprocess the LLP dataset as the first full multi-source audio-to-image generation benchmark. The experiments are conducted on multi-source, mixed-source, and single-source audio-to-image generation tasks. The proposed MACS outperforms the current state-of-the-art methods in 17 out of the 21 evaluation indexes on all tasks and delivers superior visual quality.

音频生成跨模态多源处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。