arXiv:2503.00733eess.AScs.CL2025-03ICLR被引 8

统一预训练框架让语音模型同时胜任识别与生成任务

UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation

  • 设计编码器-解码器结构,联合训练语音表征与生成能力
  • 在语音识别、语音合成等任务上表现媲美专用模型
  • 适合需要通用语音基础模型的研究者和开发者

预训练与表征学习在现代语音处理中日益重要。然而,不同应用仍依赖不同的基础模型,因主流预训练方法仅针对判别或生成任务设计。本文首次尝试构建统一的语音预训练框架,涵盖两类任务。通过合理的预训练设计,我们联合学习一个表征编码器与生成音频解码器,可应用于多种场景。提出UniWav框架,其在语音识别、文本到语音、语音分词任务上表现媲美各自专用的基础模型。结果表明,可通过单一通用语音基础模型替代多个专用模型,降低预训练成本与开销。

原文摘要 · Abstract (English)

Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training techniques are either designed for discriminative tasks or generative tasks. In this work, we make the first attempt at building a unified pre-training framework for both types of tasks in speech. We show that with the appropriate design choices for pre-training, one can jointly learn a representation encoder and generative audio decoder that can be applied to both types of tasks. We propose UniWav, an encoder-decoder framework designed to unify pre-training representation learning and generative tasks. On speech recognition, text-to-speech, and speech tokenization, UniWav achieves comparable performance to different existing foundation models, each trained on a specific task. Our findings suggest that a single general-purpose foundation model for speech can be built to replace different foundation models, reducing the overhead and cost of pre-training.

语音生成统一模型预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。