用高保真声学模拟数据训练,显著提升多通道语音增强效果。
Improving multichannel speech enhancement through accurate room-acoustic simulations
- 采用波基与几何结合的高保真声学模拟生成训练数据
- 相比低保真方法,词错误率中位数降低38%
- 适合关注语音增强数据质量的研究者
房间声学模拟广泛用于深度学习语音增强模型的训练数据扩充。现有大多数流程依赖简化的几何声学方法,而波基方法能提供更高的物理准确性。本文研究模拟精度对多通道语音增强性能的影响:在不同声学模拟方法生成的数据集上训练SpatialNet模型,并在实测数据上评估其表现。对比基于几何声学的低保真数据集与采用先进声学建模的高保真数据集,以及波基与几何结合的混合模拟数据集。结果表明,在高保真数据集上训练的模型,在中位词错误率上相较低保真方案最多降低38%。这证明高保真声学模拟数据的增广可直接提升多通道语音增强性能。
原文摘要 · Abstract (English)
Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement. While most pipelines rely on simplified geometrical acoustics, wave-based approaches offer greater physical accuracy. In this work, we examine how simulation fidelity affects multichannel speech enhancement performance. To this end, we train SpatialNet on datasets augmented with different room-acoustic simulation methods and evaluate the resulting models on measured data. We compare lower-fidelity datasets based on geometrical acoustics with a high-fidelity dataset using advanced acoustic modelling and a hybrid combination of wave-based and geometrical acoustics simulations. Training on the high-fidelity dataset results in an up to 38 % relative reduction in median word error rate compared to the lower-fidelity alternatives. These results show that augmentation with high-fidelity room-acoustic simulations directly translates into improved multichannel speech enhancement performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。