解决语音匿名化中跨语言性能下降问题,提升日语和中文可用性。
Mitigating Language Mismatch in SSL-Based Speaker Anonymization
- 用目标语言微调自监督语音模型,增强跨语言适配能力。
- 日语微调后语音可懂度提升,同时保护说话人隐私。
- 多语言预训练模型使系统支持更多语言,适合多语种场景。
语音匿名化旨在保护说话人身份的同时保留内容信息与语音可懂度。然而,当前多数语音匿名系统(SAS)仅在英语上开发和评估,导致其他语言性能显著下降。本文研究了日语和汉语语音中的语言不匹配问题。首先,使用日语语音微调基于自监督学习(SSL)的内容编码器,验证语言适应的有效性;随后,提出用日语微调多语言SSL模型,并在日语和汉语上评估系统性能。下游实验表明,仅用英语训练的SSL模型通过目标语言微调后,可提升语音可懂度且保持隐私性;而多语言SSL模型进一步扩展了系统的跨语言适用性。结果强调了语言适应与多语言预训练在构建鲁棒多语言语音匿名系统中的重要性。
原文摘要 · Abstract (English)
Speaker anonymization aims to protect speaker identity while preserving content information and the intelligibility of speech. However, most speaker anonymization systems (SASs) are developed and evaluated using only English, resulting in degraded utility for other languages. This paper investigates language mismatch in SASs for Japanese and Mandarin speech. First, we fine-tune a self-supervised learning (SSL)-based content encoder with Japanese speech to verify effective language adaptation. Then, we propose fine-tuning a multilingual SSL model with Japanese speech and evaluating the SAS in Japanese and Mandarin. Downstream experiments show that fine-tuning an English-only SSL model with the target language enhances intelligibility while maintaining privacy and that multilingual SSL further extends SASs' utility across different languages. These findings highlight the importance of language adaptation and multilingual pre-training of SSLs for robust multilingual speaker anonymization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。