arXiv:2603.23048cs.SDcs.AI2026-03

让语音预训练模型适应多种采样率,无需重采样

MSR-HuBERT: Self-supervised Pre-training for Adaptation to Multiple Sampling Rates

  • 用多采样率自适应下采样网络替代单采样率模块
  • 在16-48kHz范围内语音识别与全频段重建性能更优
  • 兼容原有HuBERT架构,适合语音处理研究者使用

自监督学习(SSL)推动了语音处理的发展,但现有方法通常假设单一采样率,在混合采样率数据上因时间分辨率不匹配而表现不佳。为此,我们提出MSRHuBERT,一种支持多采样率的预训练方法。基于HuBERT,将单采样率下采样卷积神经网络替换为多采样率自适应下采样网络,可将不同采样率的原始波形映射到统一时间分辨率,无需重采样。该设计实现统一的混合采样率预训练与微调。在16至48 kHz范围内的实验表明,MSRHuBERT在语音识别和全频段语音重建任务中均优于HuBERT,同时保留高频细节并建模低频语义结构。此外,MSRHuBERT保持了HuBERT的掩码预测目标和Transformer编码器,因此适用于已针对HuBERT开发的分析与改进。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has advanced speech processing. However, existing speech SSL methods typically assume a single sampling rate and struggle with mixed-rate data due to temporal resolution mismatch. To address this limitation, we propose MSRHuBERT, a multi-sampling-rate adaptive pre-training method. Building on HuBERT, we replace its single-rate downsampling CNN with a multi-sampling-rate adaptive downsampling CNN that maps raw waveforms from different sampling rates to a shared temporal resolution without resampling. This design enables unified mixed-rate pre-training and fine-tuning. In experiments spanning 16 to 48 kHz, MSRHuBERT outperforms HuBERT on speech recognition and full-band speech reconstruction, preserving high-frequency detail while modeling low-frequency semantic structure. Moreover, MSRHuBERT retains HuBERT's mask-prediction objective and Transformer encoder, so existing analyses and improvements that were developed for HuBERT can apply directly.

自监督学习语音处理多采样率HuBERT

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。