用神经声学场与检索增强预训练,生成更真实的混响数据提升语音定位精度。
Data Augmentation Using Neural Acoustic Fields With Retrieval-Augmented Pre-training
- 基于房间几何信息预训练神经声学场,实现声学环境建模。
- 通过检索外部数据集获取几何信息,适配目标房间并生成新混响数据。
- 生成的混响数据用于训练说话人距离估计模型,提升任务性能。
本文介绍梅尔研究所提交的房间脉冲响应(RIR)估计系统,参与ICASSP 2025生成式数据增强研讨会的任务1(扩充RIR数据)与任务2(提升说话人距离估计)。首先,在外部大规模数据集上对条件于房间几何的神经声学场进行预训练,该数据集包含成对的RIR与几何信息。随后,利用注册数据对神经声学场进行微调,根据目标房间是否提供几何信息,分别采用原始几何或从外部数据集中检索的几何。最后,针对任务1中指定的声源与接收器位置对,预测对应的RIR,并使用这些生成的RIR数据训练任务2中的说话人距离估计模型。
原文摘要 · Abstract (English)
This report details MERL's system for room impulse response (RIR) estimation submitted to the Generative Data Augmentation Workshop at ICASSP 2025 for Augmenting RIR Data (Task 1) and Improving Speaker Distance Estimation (Task 2). We first pre-train a neural acoustic field conditioned by room geometry on an external large-scale dataset in which pairs of RIRs and the geometries are provided. The neural acoustic field is then adapted to each target room by using the enrollment data, where we leverage either the provided room geometries or geometries retrieved from the external dataset, depending on availability. Lastly, we predict the RIRs for each pair of source and receiver locations specified by Task 1, and use these RIRs to train the speaker distance estimation model in Task 2.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。