arXiv:2511.16639eess.AScs.CL2025-11被引 3

用离散语音编码单元做自监督语音表征,省存储快训练

Codec2Vec: Self-Supervised Speech Representation Learning Using Neural Speech Codecs

  • 直接用神经语音编码的离散单元作为输入,跳过连续信号处理
  • 在SUPERB上表现媲美连续模型,存储减少16.5倍,训练快2.3倍
  • 适合资源受限场景,尤其关注效率与隐私的语音任务

近年来,神经音频编码器不仅实现了更优的音频压缩,还提升了语音合成效果。研究者正探索其作为通用声学特征提取器在各类语音处理任务中的潜力。基于此趋势,我们提出Codec2Vec,首个完全依赖离散音频编码单元的语音表征学习框架。该方法在数据存储与传输效率、训练速度及数据隐私方面具有优势。通过多种掩码预测目标推导策略,深入评估了该框架的有效性。在SUPERB基准测试中,Codec2Vec表现接近连续输入模型,同时存储需求降低最多16.5倍,训练时间缩短2.3倍,展现出良好的可扩展性与高效性。

原文摘要 · Abstract (English)

Recent advancements in neural audio codecs have not only enabled superior audio compression but also enhanced speech synthesis techniques. Researchers are now exploring their potential as universal acoustic feature extractors for a broader range of speech processing tasks. Building on this trend, we introduce Codec2Vec, the first speech representation learning framework that relies exclusively on discrete audio codec units. This approach offers several advantages, including improved data storage and transmission efficiency, faster training, and enhanced data privacy. We explore masked prediction with various training target derivation strategies to thoroughly understand the effectiveness of this framework. Evaluated on the SUPERB benchmark, Codec2Vec achieves competitive performance compared to continuous-input models while reducing storage requirements by up to 16.5x and training time by 2.3x, showcasing its scalability and efficiency.

语音表征自监督编码器高效学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。