arXiv:2510.17662cs.SDcs.CL2025-10被引 1

让语音模型学会识别说话人,提升验证与分析效果

DELULU: Discriminative Embedding Learning Using Latent Units for Speaker-Aware Self-Trained Speech Foundational Model

  • 用语音验证模型的帧级特征引导聚类,引入说话人区分性先验
  • 说话人验证EER降低62%,零样本分析任务全面领先
  • 无需微调即可通用,适合说话人相关语音任务

自监督语音模型在内容驱动任务上表现优异,但在说话人区分特征捕捉方面仍有不足,制约其在验证、分割和画像等应用中的表现。本文提出 extsc{DELULU},一种面向说话人的自训练基础模型,通过将说话人信息融入伪标签生成过程来弥补这一缺陷。该模型利用前沿说话人验证模型ReDimNet的帧级嵌入,指导预训练阶段的k-means聚类,引入说话人区分性归纳偏置,使表征学习更贴近说话人身份。实验表明,DELULU在多种说话人相关任务中显著优于现有自监督模型,说话人验证的等错误率(EER)相对提升最高达62%,在零样本画像任务(包括性别、年龄、口音、说话人计数)上也取得一致提升,甚至超越其教师模型在零样本评估的表现。结果证明,DELULU是强大的通用说话人感知编码器,无需任务微调即可实现优异性能。

原文摘要 · Abstract (English)

Self-supervised speech models have achieved remarkable success on content-driven tasks, yet they remain limited in capturing speaker-discriminative features critical for verification, diarization, and profiling applications. We introduce \textsc{DELULU}, a speaker-aware self-trained foundational model that addresses this limitation by incorporating speaker-informed structure into pseudo-label generation. DELULU leverages frame-level embeddings from ReDimNet, a state-of-the-art speaker verification model, to guide k-means clustering during pre-training, introducing a speaker-discriminative inductive bias that aligns representation learning with speaker identity. DELULU significantly outperforms prior SSL models across a range of speaker-centric tasks, achieving up to \textbf{62\% relative improvement} in equal error rate (EER) for speaker verification and consistent gains on zero-shot profiling tasks including gender, age, accent, and speaker counting; notably surpassing even its teacher model on zero-shot evaluations. Our findings demonstrate that \textbf{DELULU is a strong universal encoder for speaker-aware speech processing}, enabling superior performance without task-specific fine-tuning.

说话人识别自监督学习语音模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。