arXiv:2512.15224eess.AS2025-12中稿 · ASRU25被引 4

评估自监督语音模型在说话人分离与辨识中的表现,发现现有评测存在数据集单一、系统多样性不足问题。

On the Use of Self-Supervised Representation Learning for Speaker Diarization and Separation

  • 使用wav2vec2.0和WavLM等自监督模型提取语音表征
  • 发现现有评测数据集缺乏多样性,限制了模型性能评估
  • 适合关注语音识别下游任务的开发者与研究者

近年来,wav2vec2.0和WavLM等自监督语音模型在多个下游语音任务中显著提升了性能,尤其在低资源场景下。然而,针对说话人辨识与语音分离等与说话人身份相关的任务,相关评估仍较为有限。本文系统考察了近期自监督语音表征在这两项任务上的表现,揭示了当前文献中存在的关键缺陷:现有基准测试存在数据集多样性不足、下游系统类型单一等问题,影响了对模型真实能力的全面评估。

原文摘要 · Abstract (English)

Self-supervised speech models such as wav2vec2.0 and WavLM have been shown to significantly improve the performance of many downstream speech tasks, especially in low-resource settings, over the past few years. Despite this, evaluations on tasks such as Speaker Diarization and Speech Separation remain limited. This paper investigates the quality of recent self-supervised speech representations on these two speaker identity-related tasks, highlighting gaps in the current literature that stem from limitations in the existing benchmarks, particularly the lack of diversity in evaluation datasets and variety in downstream systems associated to both diarization and separation.

自监督学习说话人辨识语音分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。