用更真实的混响数据训练语音增强模型,能显著提升实际场景表现。
Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement

- 用波-几何混合仿真生成高保真混响数据替代传统图像源法。
- 在未见实测环境上,语音识别错误率降低21.3%,客观指标更优。
- 适合做语音增强数据构建或追求真实场景泛化能力的研究者。
我们研究了合成混响脉冲响应(RIR)数据的真实性对DeepFilterNet3单通道语音增强模型训练的影响。对比了采用图像源法(ISM)生成的DNS4 RIR数据集与通过波-几何声学混合仿真生成的更高保真度数据集。不拆解单一仿真因素,而是比较完整的RIR生成流程,同时保持增强模型不变。模型在未见过的实测RIR上使用客观语音增强指标和下游自动语音识别(ASR)进行评估。结果显示,使用高保真数据训练的模型在客观指标上略有提升,在未见环境中显著降低ASR词错误率(平均降低21.3%)。尽管未归因于具体建模组件,但表明提升合成声学训练数据的整体真实性,有助于提高DeepFilterNet3对未知测量环境的泛化能力。
原文摘要 · Abstract (English)
We investigate how the realism of synthetic room impulse response (RIR) datasets affects the training of DeepFilterNet3 for single-channel speech enhancement. We compare a DNS4 image-source-method (ISM) RIR dataset with a higher-acoustic-fidelity dataset generated using hybrid wave-based and geometrical acoustics simulation. Rather than isolating individual simulation factors, we compare complete RIR generation pipelines while keeping the enhancement model unchanged. Models are evaluated on unseen measured RIRs using objective speech enhancement metrics and downstream automatic speech recognition (ASR). Training with the higher-fidelity dataset consistently yields modest improvements in objective metrics and substantially lower ASR word error rates than the ISM dataset. Although the experiments do not attribute these gains to individual modelling components, they show that increasing the overall realism of synthetic acoustic training data improves the generalization of DeepFilterNet3 to unseen measured environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。