用检索增强生成提升文本转音频的零样本和少样本表现
Audiobox TTA-RAG: Improving Zero-Shot and Few-Shot Text-To-Audio with Retrieval-Augmented Generation
- 通过检索相似音频并融合到文本条件中,增强生成能力
- 在多个指标上显著提升零样本与少样本性能,效果明显
- 无需标注数据库,更适用于真实场景,适合音频生成研究者
本文聚焦于提升文本转音频(TTA)在零样本和少样本设置下的性能(即生成未见过或罕见的音频事件)。受大语言模型中检索增强生成(RAG)成功的启发,我们提出基于Audiobox(一种流匹配音频生成模型)的新型检索增强型TTA方法——Audiobox TTA-RAG。不同于仅依赖文本条件的原始Audiobox TTA方案,我们通过在条件输入中同时加入文本和检索到的音频样本来扩展生成过程。该检索方法无需外部数据库包含标注音频,具备更强实用性。实验表明,所提模型能有效利用检索到的音频样本,在多个评估指标上显著提升零样本和少样本TTA性能,且保持对领域内音频的语义一致性生成能力。
原文摘要 · Abstract (English)
This work focuses on improving Text-To-Audio (TTA) generation on zero-shot and few-shot settings (i.e. generating unseen or uncommon audio events). Inspired by the success of Retrieval-Augmented Generation (RAG) in Large Language Models, we propose Audiobox TTA-RAG, a novel retrieval-augmented TTA approach based on Audiobox, a flow-matching audio generation model. Unlike the vanilla Audiobox TTA solution that generates audio conditioned on text only, we extend the TTA process by augmenting the conditioning input with both text and retrieved audio samples. Our retrieval method does not require the external database to have labeled audio, offering more practical use cases. We show that the proposed model can effectively leverage the retrieved audio samples and significantly improve zero-shot and few-shot TTA performance, with large margins on multiple evaluation metrics, while maintaining the ability to generate semantically aligned audio for the in-domain setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。