ViP-VL用向量量化学习提升越南语语音预训练效率与性能。
ViP-VL: Vietnamese Self-supervised Speech Pretraining Model with Vector-Quantization Learning
- 通过声学堆叠与感受野对齐实现8倍下采样同步处理。
- 在1.7万小时无标签越南语数据上预训练,四项任务达新最佳。
- 适合越南语语音识别、情感分析等方向研究者使用。
我们提出ViP-VL,一种基于向量量化学习的高效越南语自监督语音预训练模型。为弥合高分辨率音频与高效处理之间的差距,ViP-VL在ChunkFormer架构中引入声学堆叠与感受野对齐,实现同步8倍下采样率,同时在BEST-RQ框架预训练过程中采用专用掩码选择策略,进一步增强表示鲁棒性。模型在17,000小时未标注越南语语音数据上进行预训练,在自动语音识别、语音情感识别、方言分类和说话人验证四个主要下游任务上均取得新最优结果。为促进未来越南语语音技术研究与开发,我们已将预训练权重与实现代码公开于github.com/khanld/chunkformer。
原文摘要 · Abstract (English)
We present ViP-VL, an efficient Vietnamese Self-supervised speech Pretraining model leveraging Vector-quantization Learning. To bridge the gap between high-resolution audio and efficient processing, ViP-VL incorporates Acoustic Stacking and Receptive Field Alignment to enable a synchronized 8x subsampling rate within the ChunkFormer architecture, while further enhancing representation robustness through a specialized Mask Selection Strategy during pretraining on the BEST-RQ framework. Pretrained on 17,000 hours of unlabeled Vietnamese speech, our model establishes new state-of-the-art results across four major downstream tasks: Automatic Speech Recognition, Speech Emotion Recognition, Dialect Classification, and Speaker Verification. To facilitate future research and the development of high-performance Vietnamese speech technologies, we publicly release our pretrained weights and implementation at github.com/khanld/chunkformer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。