构建高质量音图数据集,提升音频生成图像的表达力与对齐精度。
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework

- 构建32.3万对音图文本三模态数据集,支持精准音频到图像生成。
- 提出AudioCanvas模型,在新数据集上实现更逼真且对齐的生成效果。
- 适合研究跨模态生成、音画联动或想用强文本图像模型改进音频生成者。
作为跨模态生成的重要分支,从音频生成静态视觉内容(即音频到图像,A2I)近年来受到越来越多关注。尽管现代文本到图像(T2I)模型在视觉质量上表现优异,但传统A2I数据集普遍缺乏高保真图像和精确的跨模态对齐,导致现有方法难以通过微调强大T2I模型实现高质量生成,限制了实际应用。为填补这一空白,我们提出了A2I-Set,一个统一的高质量三模态数据集,包含32.3万对配对的音频、图像及详细文本描述,专为音频视觉研究设计,涵盖音频条件图像生成任务。此外,我们通过人工监督构建了新的混合源测试集。我们进一步提出基于该数据集微调的A2I模型AudioCanvas。实验表明,AudioCanvas在视觉表现力和跨模态对齐方面均优于现有方法。数据集与代码已开源:https://github.com/gdx012/A2I-Generation。
原文摘要 · Abstract (English)
As an important subfield of cross-modal generation, synthesizing static visual content in the form of images from audio, namely audio-to-image (A2I) generation, has attracted increasing research attention in recent years. Nevertheless, despite the remarkable visual quality of modern text-to-image (T2I) models, the performance of A2I remains fundamentally limited by traditional datasets, which often lack both high-fidelity images and precise cross-modal alignment. As a result, existing methods still struggle to achieve high-quality audio-to-image generation through finetuning strong T2I models, thereby constraining practical applications in this area. Motivated by this gap, we introduce A2I-Set, a unified, high-quality tri-modal dataset consisting of 323K paired audio, images, and detailed text captions, specifically designed for audio-visual research, including audio-conditioned image generation. Besides, we developed a new mixed-source test set for the A2I task through human supervision. We further propose an A2I model, AudioCanvas, fine-tuned on our A2I-Set. Experiments show that AudioCanvas achieves more visually expressive as well as cross-modal alignment results that generally outperforming existing approaches. Our dataset and source code are available at https://github.com/gdx012/A2I-Generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。