arXiv:2606.17967cs.CL2026-06中稿 · Interspeech 2026

让语音大模型学会分离说话人和内容信息,提升下游任务效果

Learning task-specific subspaces via interventional post-training of speech foundation models

  • 用干预数据+对比学习,把语音表征拆成说话人和内容两个独立子空间
  • 在跨域说话人验证中准确率提升,且子空间间信息互不干扰
  • 适合需要解耦语音特征的场景,如隐私保护、个性化语音系统

语音基础模型在大量无标签语音数据上预训练,生成通用表征,适用于多种任务。然而,这些表征以分布式方式编码了关键语音变量,而下游任务仅依赖其中部分变化。本文提出一种后训练优化方法,采用干预式对比学习,利用干预数据集和多部分对比损失,将语音基础模型的纠缠表征空间转换为独立的内容与说话人子空间。我们在说话人验证和关键词检测任务上评估了所学表征,结果表明跨域说话人验证性能提升,并证实了学习后的子空间中说话人与内容信息实现了有效分离。

原文摘要 · Abstract (English)

Speech foundation models, pre-trained on large corpora of unlabelled speech data, produce general-purpose representations which are useful across tasks. However, these representations encode information about salient speech variables in a distributed manner, while downstream speech tasks rely on only some of this variability. In this work, we propose a post-training refinement approach using interventional contrastive learning. By leveraging an interventional dataset and multi-part contrastive loss, we learn a transformation from the entangled representation space of speech foundation models into separate content and speaker subspaces. We evaluate the learnt representations on speaker verification and keyword spotting tasks, showing improved out-of-domain speaker verification performance and evidence that speaker and content information are separated across the learned subspaces.

语音表征表征解耦对比学习说话人分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。