用模拟混语数据训练,实现多语言语音分说话人端到端识别
SAGE-LD: Towards Scalable and Generalizable End-to-End Language Diarization via Simulated Data Augmentation
- 基于可学习查询的多语言感知架构,结合模拟混语数据预训练
- 在多个基准上达到顶尖性能,相对提升23%至52%
- 适合需要跨语言、跨场景语音分离的研究与应用
本文提出一种神经语音语言分说话人模型,可在单一框架内支持任意语言组合。方法融合基于可学习查询的多语言感知架构与大规模模拟混语数据预训练,有效克服传统方法在数据稀缺和架构优化上的局限,在多样化真实环境中表现出强泛化能力。实验表明,该方法在多个语言分说话人基准上取得当前最优性能,相较此前方法相对提升23%至52%。本工作不仅推动语言分说话人研究进展,也为混语语音技术奠定基础框架。
原文摘要 · Abstract (English)
In this paper, we present a neural spoken language diarization model that supports an unconstrained span of languages within a single framework. Our approach integrates a learnable query-based architecture grounded in multilingual awareness, with large-scale pretraining on simulated code-switching data. By jointly leveraging these two components, our method overcomes the limitations of conventional approaches in data scarcity and architecture optimization, and generalizes effectively to real-world multilingual settings across diverse environments. Experimental results demonstrate that our approach achieves state-of-the-art performance on several language diarization benchmarks, with a relative performance improvement of 23% to 52% over previous methods. We believe that this work not only advances research in language diarization but also establishes a foundational framework for code-switching speech technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。