用语音转换提升阿拉伯语方言识别的跨领域鲁棒性
Voice Conversion Improves Cross-Domain Robustness for Spoken Arabic Dialect Identification
- 通过语音转换增强数据多样性,缓解方言识别中的说话人偏差
- 在四个真实场景下跨域测试,准确率最高提升34.1%
- 适合开发面向多方言的包容性语音技术的研究者
阿拉伯语方言识别(ADI)系统对大规模语音技术开发至关重要。然而,现有系统在跨领域场景下泛化能力差。本文提出基于语音转换的训练方法,显著提升模型鲁棒性,在新收集的涵盖四个真实领域的测试集上,跨域准确率最高提升34.1%。分析表明,该方法有效缓解了数据集中说话人偏差问题。研究团队公开了鲁棒的ADI模型与跨域评估数据集,以支持更具包容性的阿拉伯语语音技术研发。
原文摘要 · Abstract (English)
Arabic dialect identification (ADI) systems are essential for large-scale data collection pipelines that enable the development of inclusive speech technologies for Arabic language varieties. However, the reliability of current ADI systems is limited by poor generalization to out-of-domain speech. In this paper, we present an effective approach based on voice conversion for training ADI models that achieves state-of-the-art performance and significantly improves robustness in cross-domain scenarios. Evaluated on a newly collected real-world test set spanning four different domains, our approach yields consistent improvements of up to +34.1% in accuracy across domains. Furthermore, we present an analysis of our approach and demonstrate that voice conversion helps mitigate the speaker bias in the ADI dataset. We release our robust ADI model and cross-domain evaluation dataset to support the development of inclusive speech technologies for Arabic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。