MERL通过高质量数据准备实现真实场景语音提取领先表现
Technical Report for MERL's Real-TSE Challenge Submission
- 分四阶段训练:先合成数据预训练,再用远场真实录音微调
- 在真实远场噪声混响场景中获第二赛道第一名
- 发现DNSMOS和说话人相似性易被过优化,需警惕指标误导
目标语音提取(TSE)主要依赖在合成全重叠数据上训练的神经网络方法。真实语音提取挑战(Real-TSE Challenge)旨在提升真实世界远场、噪声和混响环境下的性能。本文描述了梅尔研究所(MERL)在该挑战中的提交方案。我们未提出新模型架构,而是聚焦于数据准备与清洗。系统采用四阶段训练:首先在全重叠混合信号及加入噪声与混响的多说话人对话数据上进行预训练;随后利用带伪目标的真实远场录音(目标来自近场麦克风处理信号)进行适应。提交结果在第二赛道排名第一,凸显高质量数据准备的关键作用。此外,我们观察到DNSMOS和说话人相似性指标易受过优化影响,因此通过对抗攻击测试其鲁棒性。结果显示,这些指标可被推至极端值,而词错误率(WER)和基于语音活动检测的F1分数保持稳定。
原文摘要 · Abstract (English)
Target speech extraction (TSE) has largely been dominated by neural network-based approaches trained and evaluated on synthetic fully overlapped data. The Real-TSE Challenge aims to advance performance on real-world far-field noisy and reverberant recordings. This technical report describes MERL's submission to the Real-TSE Challenge. Rather than proposing a novel model architecture, we built upon the baseline model and focused primarily on data preparation and cleaning. Our system was trained in four stages, beginning with pre-training on fully overlapped mixtures and simulated multi-talker conversations with noise and reverberation applied to both the mixture and the enrollment utterances. We then adapted the model to real-world conditions using noisy far-field recordings with pseudo-targets derived from processed close-talk microphone signals. Our submission achieved first place in the second track, demonstrating the critical importance of high-quality data preparation. Furthermore, we observed that DNSMOS and speaker similarity are susceptible to over-optimization, motivating an investigation of their robustness using adversarial attacks. The results show that both metrics can be driven to extreme values without degrading the token error rate or the VAD-based F1 score.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。