arXiv:2603.24432cs.SDcs.CL2026-03

根据样本难易度动态调整学习策略,提升大规模语音验证准确率。

What and When to Learn: CURriculum Ranking Loss for Large-Scale Speaker Verification

  • 用子中心余弦置信度在线评估样本难易,分三类梯度更新。
  • 在VoxCeleb1-O上将错误率降低86.8%,SITW上降60.0%。
  • 适合处理标注不全或数据质量差的大规模语音验证任务。

大规模语音验证仍面临挑战,因固定间隔损失对所有样本一视同仁。我们假设错误标注或退化样本会引入噪声梯度,破坏紧凑的说话人流形。提出Curry(CURriculum Ranking)自适应损失,通过子中心ArcFace的置信度得分,利用运行批次统计在线评估样本难度,将样本分为易、中、难三类,无需额外标注。可学习权重引导模型从稳定身份基础,经流形优化,到边界锐化。据我们所知,这是迄今训练规模最大的语音验证系统。在VoxCeleb1-O和SITW上,Curry相较于子中心ArcFace基线分别将错误率降低86.8%和60.0%,为不完美大规模数据上的鲁棒语音验证建立新范式。

原文摘要 · Abstract (English)

Speaker verification at large scale remains an open challenge as fixed-margin losses treat all samples equally regardless of quality. We hypothesize that mislabeled or degraded samples introduce noisy gradients that disrupt compact speaker manifolds. We propose Curry (CURriculum Ranking), an adaptive loss that estimates sample difficulty online via Sub-center ArcFace: confidence scores from dominant sub-center cosine similarity rank samples into easy, medium, and hard tiers using running batch statistics, without auxiliary annotations. Learnable weights guide the model from stable identity foundations through manifold refinement to boundary sharpening. To our knowledge, this is the largest-scale speaker verification system trained to date. Evaluated on VoxCeleb1-O, and SITW, Curry reduces EER by 86.8\% and 60.0\% over the Sub-center ArcFace baseline, establishing a new paradigm for robust speaker verification on imperfect large-scale data.

语音验证自适应学习大规模训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。