用文本生成音频模型创造逼真房间回声,少数据也能行。
Adapting a Text-to-Audio Model for Room Impulse Response Generation
- 用视觉语言模型从图像-回声数据中提取声学描述,解决无配对数据难题。
- 仅用少量训练数据,生成的回声在主观测试中表现可信。
- 支持自由提示输入,适合音频增强与虚拟音效场景使用。
房间脉冲响应(RIR)能实现逼真的声学模拟,广泛应用于多媒体制作和语音数据增强。然而,获取高质量的真实世界RIR成本高昂,数据稀缺仍是数据驱动生成方法的主要挑战。本文提出一种新方法,通过适配预训练的文本到音频模型进行RIR生成,首次证明大规模生成式音频先验可有效用于此任务。为解决缺乏文本-RIR配对数据的问题,我们利用视觉语言模型从现有图像-RIR数据集中提取声学描述,构建标签管道。引入上下文学习策略,以支持推理时的自由形式用户提示。主观听感测试表明,该模型在使用较少训练数据的情况下,仍能生成可信的RIR。音频示例可在演示网站上查看。
原文摘要 · Abstract (English)
Room Impulse Responses (RIRs) enable realistic acoustic simulation, with applications ranging from multimedia production to speech data augmentation. However, acquiring high-quality real-world RIRs is labor-intensive, and data scarcity remains a challenge for data-driven RIR generation approaches. In this paper, we propose a novel approach to RIR generation by adapting a pre-trained text-to-audio model, demonstrating for the first time that large-scale generative audio priors can be effectively leveraged for this task. To address the lack of text-RIR paired data, we utilize a labeling pipeline leveraging vision-language models to extract acoustic descriptions from existing image-RIR datasets. We introduce an in-context learning strategy to accommodate free-form user prompts during inference. Evaluations including a subjective listening test demonstrate that our model generates plausible RIRs with substantially less training data. Audio examples are available on our demo website.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。