arXiv:2509.15579cs.CLcs.SD2025-09被引 2

提出分块自监督学习,提升语音实时与离线预训练效果。

Chunk Based Speech Pre-training with High Resolution Finite Scalar Quantization

  • 分块处理语音,利用前后块信息恢复被掩码帧
  • 高分辨率量化码本(数百万词表)增强下游任务迁移能力
  • 适用于实时语音交互场景,兼顾低延迟与高性能

随着语音技术快速发展,低延迟语音人机交互日益重要。自监督学习是推动该领域进步的关键因素,但现有方法多基于完整语句假设,在流式应用中常需妥协。本文提出分块自监督学习(Chunk SSL),统一支持流式与离线语音预训练。Chunk SSL 采用掩码预测损失,鼓励声学编码器利用同一分块及前序分块的未掩码帧恢复被掩码语音帧的索引。提出复制拼接数据增强策略以高效实现分块预训练。采用有限标量量化(FSQ)模块对输入语音特征进行离散化,研究表明高分辨率FSQ码本(词汇量达数百万)有助于知识迁移。预训练阶段使用组掩码预测损失,缓解大码本带来的高内存与计算开销。在Librispeech和Must-C数据集上的语音识别与语音翻译任务实验表明,该方法在流式与离线模式下均达到具有竞争力的性能。

原文摘要 · Abstract (English)

Low latency speech human-machine communication is becoming increasingly necessary as speech technology advances quickly in the last decade. One of the primary factors behind the advancement of speech technology is self-supervised learning. Most self-supervised learning algorithms are designed with full utterance assumption and compromises have to made if partial utterances are presented, which are common in the streaming applications. In this work, we propose a chunk based self-supervised learning (Chunk SSL) algorithm as an unified solution for both streaming and offline speech pre-training. Chunk SSL is optimized with the masked prediction loss and an acoustic encoder is encouraged to restore indices of those masked speech frames with help from unmasked frames in the same chunk and preceding chunks. A copy and append data augmentation approach is proposed to conduct efficient chunk based pre-training. Chunk SSL utilizes a finite scalar quantization (FSQ) module to discretize input speech features and our study shows a high resolution FSQ codebook, i.e., a codebook with vocabulary size up to a few millions, is beneficial to transfer knowledge from the pre-training task to the downstream tasks. A group masked prediction loss is employed during pre-training to alleviate the high memory and computation cost introduced by the large codebook. The proposed approach is examined in two speech to text tasks, i.e., speech recognition and speech translation. Experimental results on the \textsc{Librispeech} and \textsc{Must-C} datasets show that the proposed method could achieve very competitive results for speech to text tasks at both streaming and offline modes.

自监督学习语音预训练分块处理高分辨率量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。