通过全局局部联合分类提升说话人区分能力
Joint Learning Global-Local Speaker Classification to Enhance End-to-End Speaker Diarization and Recognition
- 用聚类说话人作全局标签,重编码内部说话人作局部标签
- 在三个数据集上达到优于或相当现有方法的性能
- 无需大量真实对话数据,适合端到端语音系统优化
大音频语言模型在端到端说话人分离与识别中表现优异,但受限于大规模对话数据稀缺及缺乏显式说话人表征优化,其说话人区分能力仍有限。为此,我们提出GLSC-SDR框架,联合训练说话人分类与分离识别任务。引入全局-局部说话人分类策略:以聚类说话人作为全局标签,重新编码簇内说话人作为局部标签。该分层设计增强细粒度说话人区分能力,同时保持语义转录准确性。在AliMeeting、AISHELL-4和AMI-SDM数据集上的实验表明,GLSC-SDR在不依赖大规模真实对话数据的前提下,性能优于或相当于基于模拟和多编码器的方法。
原文摘要 · Abstract (English)
Large Audio-Language Models (LALMs) have demonstrated remarkable performance in end-to-end speaker diarization and recognition. However, their speaker discriminability remains limited due to the scarcity of large-scale conversational data and the absence of explicit speaker representation optimization. To address this, we propose GLSC-SDR, a paradigm that jointly trains speaker classification with diarization and recognition. We further introduce a Global-Local Speaker Classification strategy, which uses clustered speakers as global labels and re-encoded intra-cluster speakers as local labels. This hierarchical design enhances fine-grained speaker discrimination while preserving semantic transcription accuracy. Experiments on AliMeeting, AISHELL-4, and AMI-SDM demonstrate that GLSC-SDR achieves competitive or superior performance compared to simulation-based and multi-encoder approaches, without relying on large-scale real conversational data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。