arXiv:2511.01091cs.SD2025-11

用反馈机制让大模型自动补全缺失音效,提升文本转音频质量。

AudioRAG+: Feedback-driven Retrieval-augmented Audio Generation with Large Audio Language Models

  • 基于大音频语言模型分析生成结果,主动检索缺失音效
  • 在多个模型上验证,显著提升对特定声音事件的生成能力
  • 无需额外训练,适合希望提升音效细节的研究者

我们提出一种通用的反馈驱动型检索增强生成(RAG)方法,利用大音频语言模型(LALMs)解决文本到音频(TTA)生成中特定声音事件缺失或生成不准确的问题。与以往需从头训练专用模型的RAG方法不同,本方法通过LALMs分析生成输出,从外部数据库中检索预训练模型难以生成的概念,并将其融入生成过程。实验表明,该方法不仅增强了LALMs识别缺失声音事件的能力,还在多种模型上取得性能提升,优于现有专门设计的RAG方法。

原文摘要 · Abstract (English)

We propose a general feedback-driven retrieval-augmented generation (RAG) approach that leverages Large Audio Language Models (LALMs) to address the missing or imperfect synthesis of specific sound events in text-to-audio (TTA) generation. Unlike previous RAG-based TTA methods that typically train specialized models from scratch, we utilize LALMs to analyze audio generation outputs, retrieve concepts that pre-trained models struggle to generate from an external database, and incorporate the retrieved information into the generation process. Experimental results show that our method not only enhances the ability of LALMs to identify missing sound events but also delivers improvements across different models, outperforming existing RAG-specialized approaches.

音频生成检索增强大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。