arXiv:2506.02083cs.SDcs.AI2025-06中稿 · Interspeech 2025, …被引 4

让语音识别模型自动分离说话人特征与语言信息,提升跨语言识别准确率。

LASPA: Language Agnostic Speaker Disentanglement with Prefix-Tuned Cross-Attention

  • 用前缀调优的交叉注意力机制,联合学习说话人与语言特征。
  • 在多语言数据集上等错误率降低,跨语言泛化能力强。
  • 适合需要跨语言语音识别的应用场景,如国际客服系统。

说话人识别模型在多语言环境下因语言信息与说话人嵌入纠缠而面临挑战。口音、声带结构与语言发音特征的重叠使得语言与说话人信息难以分离。为解决此问题,我们提出一种新型解耦学习策略,通过前缀调优的交叉注意力实现联合学习。该方法在说话人切换语言时尤为有效。实验结果表明,模型在单语和多语设置下均表现良好,包括未见过的语言。显著提升了多个数据集上的等错误率(EER),验证了其从说话人嵌入中有效分离语言信息的能力,增强了复杂语言环境下的识别性能。

原文摘要 · Abstract (English)

Speaker recognition models face challenges in multi-lingual settings due to the entanglement of linguistic information within speaker embeddings. The overlap between vocal traits such as accent, vocal anatomy, and a language's phonetic structure complicates separating linguistic and speaker information. Disentangling these components can significantly improve speaker recognition accuracy. To this end, we propose a novel disentanglement learning strategy that integrates joint learning through prefix-tuned cross-attention. This approach is particularly effective when speakers switch between languages. Experimental results show the model generalizes across monolingual and multi-lingual settings, including unseen languages. Notably, the proposed model improves the equal error rate across multiple datasets, highlighting its ability to separate language information from speaker embeddings and enhance recognition in diverse linguistic conditions.

说话人分离多语言注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。