用声音事件增强文本提示,实现多风格音频精准生成
Text Prompt is Not Enough: Sound Event Enhanced Prompt Adapter for Target Style Audio Generation
- 通过文本与参考音频的交叉注意力提取动态风格嵌入
- 在测试集上达到26.94的弗雷歇距离和1.82的KL散度
- 适合需要高保真风格迁移的音频生成研究者
当前主流音频生成方法主要依赖简单文本提示,难以捕捉多风格生成所需的细微差异。为此,本文提出声事件增强提示适配器(Sound Event Enhanced Prompt Adapter),不采用传统静态全局风格迁移,而是通过文本与参考音频之间的交叉注意力提取风格嵌入,实现自适应风格控制。随后利用自适应层归一化提升模型表达多种风格的能力。同时,构建了声事件参考风格迁移数据集(SERST),支持结合文本和音频参考的双提示音频生成。实验表明,该模型表现稳健,在弗雷歇距离(26.94)和KL散度(1.82)上优于Tango、AudioLDM和AudioGen,生成音频与参考音频高度相似。演示、代码和数据集均已公开。
原文摘要 · Abstract (English)
Current mainstream audio generation methods primarily rely on simple text prompts, often failing to capture the nuanced details necessary for multi-style audio generation. To address this limitation, the Sound Event Enhanced Prompt Adapter is proposed. Unlike traditional static global style transfer, this method extracts style embedding through cross-attention between text and reference audio for adaptive style control. Adaptive layer normalization is then utilized to enhance the model's capacity to express multiple styles. Additionally, the Sound Event Reference Style Transfer Dataset (SERST) is introduced for the proposed target style audio generation task, enabling dual-prompt audio generation using both text and audio references. Experimental results demonstrate the robustness of the model, achieving state-of-the-art Fréchet Distance of 26.94 and KL Divergence of 1.82, surpassing Tango, AudioLDM, and AudioGen. Furthermore, the generated audio shows high similarity to its corresponding audio reference. The demo, code, and dataset are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。