arXiv:2512.07168cs.SDcs.AI2025-12被引 3

用自监督框架学习低帧率语音表示,实现高效压缩与高保真重建。

JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention

  • 分两阶段:先用JEPA+DAAM在隐空间预测语义特征,不依赖波形重建
  • 生成47.5个/秒的可逆令牌,压缩率高且适配语言模型
  • 通过密度自适应门控发现语音层次结构,适合语音编码与生成任务

我们提出一种两阶段自监督框架,结合联合嵌入预测架构(JEPA)与密度自适应注意力机制(DAAM),用于学习鲁棒语音表征。第一阶段利用JEPA与DAAM在隐空间中通过掩码预测学习语义音频特征,完全脱离波形重建。第二阶段基于这些表征,采用有限标量量化(FSQ)和混合进制打包方案进行高效令牌化,并通过HiFi-GAN解码器实现高保真波形重建。通过在JEPA编码器中引入基于高斯混合的密度自适应门控,模型实现自适应时间特征选择,在2.5赫兹低帧率下发现语音的层次结构。最终生成的令牌速率为47.5个/秒,具有可逆性、高度压缩性及语言模型友好性,性能媲美甚至优于现有神经音频编解码器。

原文摘要 · Abstract (English)

We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM to learn semantic audio features via masked prediction in latent space, fully decoupled from waveform reconstruction. Stage~2 leverages these representations for efficient tokenization using Finite Scalar Quantization (FSQ) and a mixed-radix packing scheme, followed by high-fidelity waveform reconstruction with a HiFi-GAN decoder. By integrating Gaussian mixture-based density-adaptive gating into the JEPA encoder, the model performs adaptive temporal feature selection and discovers hierarchical speech structure at a low frame rate of 2.5~Hz. The resulting tokens (47.5 tokens/sec) provide a reversible, highly compressed, and language-model-friendly representation that is competitive with, and often more efficient than, existing neural audio codecs.

语音表征自监督学习音频编码令牌化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。