用大模型生成音频对应图文提示,提升音画生成一致性。
CatchPhrase: EXPrompt-Guided Encoder Adaptation for Audio-to-Image Generation
- 用LLM和音频描述模型从弱标签生成丰富语义提示。
- 通过多模态筛选选出最匹配的提示,减少音画错配。
- 轻量适配网络让现成文生图模型支持音频输入。
我们提出CatchPhrase,一种新型音频到图像生成框架,旨在缓解音频输入与生成图像间的语义错配问题。尽管多模态编码器的进步推动了跨模态生成的发展,但同音词和听觉幻觉带来的歧义仍阻碍准确对齐。为此,CatchPhrase利用大语言模型(LLMs)和音频字幕模型(ACMs),从弱类别标签中生成增强的跨模态语义提示(EXPrompt Mining)。为应对类别级和实例级错配,采用多模态过滤与检索机制,为每个音频样本选择最语义对齐的提示(EXPrompt Selector)。随后,训练一个轻量映射网络,将预训练文本到图像生成模型适配至音频输入。在多个音频分类数据集上的大量实验表明,CatchPhrase有效提升了音画对齐度,并持续改善生成质量。
原文摘要 · Abstract (English)
We propose CatchPhrase, a novel audio-to-image generation framework designed to mitigate semantic misalignment between audio inputs and generated images. While recent advances in multi-modal encoders have enabled progress in cross-modal generation, ambiguity stemming from homographs and auditory illusions continues to hinder accurate alignment. To address this issue, CatchPhrase generates enriched cross-modal semantic prompts (EXPrompt Mining) from weak class labels by leveraging large language models (LLMs) and audio captioning models (ACMs). To address both class-level and instance-level misalignment, we apply multi-modal filtering and retrieval to select the most semantically aligned prompt for each audio sample (EXPrompt Selector). A lightweight mapping network is then trained to adapt pre-trained text-to-image generation models to audio input. Extensive experiments on multiple audio classification datasets demonstrate that CatchPhrase improves audio-to-image alignment and consistently enhances generation quality by mitigating semantic misalignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。