arXiv:2505.04457cs.SDcs.CL2025-05中稿 · IEEE WASPAA2025被引 8

Miipher-2可高效修复百万小时多语言语音,用于大模型训练数据清洗。

Miipher-2: A Universal Speech Restoration Model for Million-Hour Scale Data Restoration

  • 用冻结的通用语音模型提取特征,无需文本或说话人条件。
  • 在3000小时多语种数据上训练,实现毫秒级实时处理。
  • 适合大规模语音数据清洗,支持100张消费级显卡并行处理。

训练数据清洗是生成式语音恢复(SR)的新应用。本文提出Miipher-2,一种专为百万小时级语音数据设计的语音恢复模型,用于大模型如大语言模型的训练数据清洗。关键挑战包括对未见语言的泛化能力、无需显式条件(如文本、说话人ID)运行,以及计算效率。Miipher-2采用冻结预训练的通用语音模型(USM),支持超过300种语言,作为鲁棒的无条件特征提取器。为优化效率并减少内存占用,模型引入并行适配器,从噪声输入预测干净的USM特征,并使用WaveFit神经声码器进行波形合成。上述组件在3,000小时多语言、录音室质量的录音数据上训练,且USM参数保持固定。实验结果表明,Miipher-2在所有测试语言中,词错误率、说话人相似度及客观与主观音质评分上均优于或媲美传统SR模型。该模型可在消费级加速器上高效运行,实时因子达0.0078,仅用100张此类加速器即可在约三天内完成百万小时语音数据的处理。

原文摘要 · Abstract (English)

Training data cleaning is a new application for generative model-based speech restoration (SR). This paper introduces Miipher-2, an SR model designed for million-hour scale data, for training data cleaning for large-scale generative models like large language models. Key challenges addressed include generalization to unseen languages, operation without explicit conditioning (e.g., text, speaker ID), and computational efficiency. Miipher-2 utilizes a frozen, pre-trained Universal Speech Model (USM), supporting over 300 languages, as a robust, conditioning-free feature extractor. To optimize efficiency and minimize memory, Miipher-2 incorporates parallel adapters for predicting clean USM features from noisy inputs and employs the WaveFit neural vocoder for waveform synthesis. These components were trained on 3,000 hours of multi-lingual, studio-quality recordings with augmented degradations, while USM parameters remained fixed. Experimental results demonstrate Miipher-2's superior or comparable performance to conventional SR models in word-error-rate, speaker similarity, and both objective and subjective sound quality scores across all tested languages. Miipher-2 operates efficiently on consumer-grade accelerators, achieving a real-time factor of 0.0078, enabling the processing of a million-hour speech dataset in approximately three days using only 100 such accelerators.

语音恢复数据清洗多语言高效处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。