用生成的混响数据提升语音距离估计精度,尤其在中长距离更明显。
Towards Improving Speaker Distance Estimation through Generative Impulse Response Augmentation

- 用位置信息生成混响响应,补充真实数据不足
- 中长距离误差从2.18米降至0.69米,平均降低超60%
- 适合做声学建模与语音定位相关研究者参考
ICASSP 2025的室内外声学与语音距离估计(SDE)挑战赛探讨了增强房间混响响应(RIR)数据对SDE模型性能的提升效果。本工作通过开源FastRIR生成器,仅基于说话人与听者位置生成RIR,并设计质量筛选机制确保生成数据与挑战赛真实RIR一致。采用超参数优化进行模型微调。结果表明,在GWA房间中,五种位置的平均绝对误差(MAE)由1.66米降至0.6米;在Treble房间中,由2.18米降至0.69米,验证了数据增强策略显著提升了中长距离的估计精度。
原文摘要 · Abstract (English)
The Room Acoustics and Speaker Distance Estimation (SDE) Challenge at ICASSP 2025 explores the effectiveness of augmented room impulse response (RIR) data for improving SDE model performance. This challenge at GenDARA involves generating RIRs to supplement sparse datasets and fine-tuning SDE models with the augmented data. We employ the open-source fast diffuse room impulse response generator (FastRIR) conditioned only on speaker and listener locations. We design a quality filter to ensure generated RIR alignment with challenge RIRs, and hyperparameter optimization is employed for model fine-tuning. Our approach reduces the mean absolute error (MAE) of the five positions from 1.66m to 0.6m for GWA rooms and from 2.18m to 0.69m for Treble rooms, with results demonstrating that the augmentation approach significantly improves estimation accuracy, particularly at medium to long distances.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。