用知识蒸馏让离散音频令牌更好识别说话人。
Text-Independent Speaker Verification Using Discrete Audio Tokens

- 用强教师模型指导离散音频令牌学习说话人特征。
- 在VoxCeleb上接近传统方法的识别准确率。
- 适合想用压缩音频做语音验证的研究者。
神经音频编解码器(NACs)实现了高效的音频压缩,并在语音合成等下游任务中表现优异。然而,其离散表示在自动说话人验证(ASV)任务中的性能始终落后于传统频谱特征。我们通过实验发现,说话人信息虽隐含在离散令牌中,但传统ASV训练范式未能有效利用。为此,我们提出跨特征知识蒸馏(CFKD)框架:通过让基于编解码器的学生模型模仿强基线Fbank教师模型的嵌入空间,提供结构化监督以充分挖掘令牌中的说话人信息。在VoxCeleb基准上的实验表明,CFKD显著提升了基于编解码器系统的ASV性能,使其逼近Fbank教师模型的准确率,凸显了离散音频令牌在多样化语音任务中的潜力。
原文摘要 · Abstract (English)
Neural audio codecs (NACs) enable efficient audio compression and have achieved success in downstream tasks such as speech synthesis. However, their discrete representations consistently underperform traditional spectral features in automatic speaker verification (ASV). We empirically demonstrate that speaker cues are implicitly preserved in discrete tokens but remain underutilized by conventional ASV training paradigms. To address this, we propose a Cross-Feature Knowledge Distillation (CFKD) framework. By guiding the codec-based student to mimic the embedding space of a strong Fbank-based teacher, CFKD provides structured supervision for effective utilization of speaker information in tokens. Experiments on the VoxCeleb benchmarks show that CFKD substantially improves the ASV performance of codec-based systems, allowing them to approach the accuracy of Fbank-based teacher models and highlighting the potential of discrete audio tokens for diverse speech tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。