用自监督框架学习低帧率语音表示,实现高效压缩与高保真重建。
JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
- 分两阶段:先用JEPA+DAAM在隐空间预测语义特征,不依赖波形重建
- 生成47.5个/秒的可逆令牌,压缩率高且适配语言模型
- 通过密度自适应门控发现语音层次结构,适合语音编码与生成任务
我们提出一种两阶段自监督框架,结合联合嵌入预测架构(JEPA)与密度自适应注意力机制(DAAM),用于学习鲁棒语音表征。第一阶段利用JEPA与DAAM在隐空间中通过掩码预测学习语义音频特征,完全脱离波形重建。第二阶段基于这些表征,采用有限标量量化(FSQ)和混合进制打包方案进行高效令牌化,并通过HiFi-GAN解码器实现高保真波形重建。通过在JEPA编码器中引入基于高斯混合的密度自适应门控,模型实现自适应时间特征选择,在2.5赫兹低帧率下发现语音的层次结构。最终生成的令牌速率为47.5个/秒,具有可逆性、高度压缩性及语言模型友好性,性能媲美甚至优于现有神经音频编解码器。
原文摘要 · Abstract (English)
We introduce a two-stage self-supervised framework that combines the Joint-Embedding Predictive Architecture (JEPA) with a Density Adaptive Attention Mechanism (DAAM) for learning robust speech representations. Stage~1 uses JEPA with DAAM to learn semantic audio features via masked prediction in latent space, fully decoupled from waveform reconstruction. Stage~2 leverages these representations for efficient tokenization using Finite Scalar Quantization (FSQ) and a mixed-radix packing scheme, followed by high-fidelity waveform reconstruction with a HiFi-GAN decoder. By integrating Gaussian mixture-based density-adaptive gating into the JEPA encoder, the model performs adaptive temporal feature selection and discovers hierarchical speech structure at a low frame rate of 2.5~Hz. The resulting tokens (47.5 tokens/sec) provide a reversible, highly compressed, and language-model-friendly representation that is competitive with, and often more efficient than, existing neural audio codecs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。