无需配对数据,用文本桥接音频与图像生成
SeeingSounds: Learning Audio-to-Visual Alignment via Text
- 用冻结的语言模型将音频映射到语义空间,再通过视觉语言模型对齐视觉
- 在多个基准上超越现有方法,零样本和有监督设置均达新高度
- 支持细粒度控制,音调变化可转化为具体描述词引导画面
我们提出SeeingSounds,一种轻量级、模块化的音频到图像生成框架,通过音频、语言与视觉的协同作用实现跨模态生成,无需配对音视频数据,也无需训练视觉生成模型。不同于将音频视为文本替代品或仅依赖音频到文本映射,本方法进行双重对齐:音频通过冻结的语言编码器投影至语义语言空间,并借助视觉语言模型在上下文中锚定至视觉域。该设计受认知神经科学启发,模拟人类感知中的自然跨模态关联。模型基于冻结的扩散主干网络,仅训练轻量级适配器,实现高效可扩展学习。此外,通过程序化文本提示生成实现细粒度与可解释性控制,例如音量或音高变化可转化为‘远处雷声’等描述性提示以引导视觉输出。在标准基准上的广泛实验表明,SeeingSounds在零样本与有监督设置下均优于现有方法,确立了可控音频到视觉生成的新基准。
原文摘要 · Abstract (English)
We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative models. Rather than treating audio as a substitute for text or relying solely on audio-to-text mappings, our method performs dual alignment: audio is projected into a semantic language space via a frozen language encoder, and, contextually grounded into the visual domain using a vision-language model. This approach, inspired by cognitive neuroscience, reflects the natural cross-modal associations observed in human perception. The model operates on frozen diffusion backbones and trains only lightweight adapters, enabling efficient and scalable learning. Moreover, it supports fine-grained and interpretable control through procedural text prompt generation, where audio transformations (e.g., volume or pitch shifts) translate into descriptive prompts (e.g., "a distant thunder") that guide visual outputs. Extensive experiments across standard benchmarks confirm that SeeingSounds outperforms existing methods in both zero-shot and supervised settings, establishing a new state of the art in controllable audio-to-visual generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。