跨生命周期语音说话人分离模型表现受年龄差异影响显著
Exploring Speech Foundation Models for Speaker Diarization Across Lifespan

- 统一框架下评估儿童、成人、老人对话的说话人分离性能
- 仅用成人数据训练的模型在儿童和老人对话中性能大幅下降
- 多年龄段联合训练可提升鲁棒性,适配特定年龄组效果更优
语音基础模型在多种语音任务中展现出强泛化能力,但其在说话人分离任务中对年龄相关域偏移的鲁棒性仍待深入研究。本文在统一的端到端神经说话人分离框架(EEND-VC)中,对涵盖儿童、成人及老年人对话的语音样本进行跨生命周期评估。比较了零样本跨年龄推理、多年龄联合训练与领域特定适配三种策略。结果表明,仅在成人语音上训练的模型在儿童和老年对话数据上性能显著下降。而跨不同年龄组的联合多年龄训练可在不降低成人对话性能的前提下提升模型鲁棒性,针对性地进行年龄组适配则能进一步提升性能,尤其在使用Whisper编码器时效果更佳。
原文摘要 · Abstract (English)
Speech foundation models have shown strong transferability across a wide range of speech applications. However, their robustness to age-related domain shift in speaker diarization remains underexplored. In this work, we present a cross-lifespan evaluation within a unified end-to-end neural diarization framework (EEND-VC), covering speech samples from conversations involving children, adults, and older adults. We compare models under zero-shot cross-age inference, joint multi-age training, and domain-specific adaptation. Results show substantial performance degradation when models trained on adult-specific speech are applied to child and older-adult conversational data. Moreover, joint multi-age training across different age groups improves robustness without reducing diarization performance in canonical adult conversations, while targeted age group adaptation yields further gains in diarization performance, particularly when using the Whisper encoder.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。