用大模型+音效库生成高保真音效,无需额外录音
SonicRAG : High Fidelity Sound Effects Synthesis Based on Retrival Augmented Generation
- 结合大模型与音效数据库,按需检索重组音频
- 提升音效多样性与质量,避免重复录制成本
- 适合音效设计、游戏影视等需要高效生成的场景
大型语言模型(LLMs)在自然语言处理和多模态学习中展现出强大能力,已成功应用于文本生成和语音合成,推动了多模态内容的理解与生成。在音效(SFX)生成领域,已有研究利用LLMs协调多个音频合成模型。然而,受限于标注数据稀缺及时间建模复杂性,现有技术仍难以实现高质量音效生成。为此,本文提出一种新框架,将LLMs与现有音效数据库结合,根据用户需求实现音效的检索、重组与合成。该方法显著提升了生成音效的多样性和保真度,同时无需额外录音成本,为音效设计与应用提供灵活高效的解决方案。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language processing (NLP) and multimodal learning, with successful applications in text generation and speech synthesis, enabling a deeper understanding and generation of multimodal content. In the field of sound effects (SFX) generation, LLMs have been leveraged to orchestrate multiple models for audio synthesis. However, due to the scarcity of annotated datasets, and the complexity of temproal modeling. current SFX generation techniques still fall short in achieving high-fidelity audio. To address these limitations, this paper introduces a novel framework that integrates LLMs with existing sound effect databases, allowing for the retrieval, recombination, and synthesis of audio based on user requirements. By leveraging this approach, we enhance the diversity and quality of generated sound effects while eliminating the need for additional recording costs, offering a flexible and efficient solution for sound design and application.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。