arXiv:2509.17052cs.SDeess.AS2025-09中稿 · ICASSP 2026被引 11

Sidon能快速修复多语言语音,提升合成音质。

Sidon: Fast and Robust Open-Source Multilingual Speech Restoration for Large-scale Dataset Cleansing

  • 用w2v-BERT 2.0和声码器组合,从杂乱录音中还原清晰语音
  • 在单张GPU上运行速度达实时的500倍,性能媲美谷歌内部模型
  • 适合语音数据清洗,提升多语言语音合成质量

大规模文本转语音(TTS)系统受限于高质量、多语言语音数据的稀缺。我们提出Sidon,一个快速、开源的语音恢复模型,可将野外采集的嘈杂语音转换为录音室级别的高质量语音,并支持数十种语言。Sidon由两个模块组成:基于w2v-BERT 2.0微调的特征预测器,用于去除噪声;以及训练用于从清洁特征中合成语音的声码器。其恢复效果与谷歌内部模型Miipher相当,适用于语音合成的数据清洗。此外,该模型计算效率极高,在单张GPU上运行速度可达实时的500倍。我们进一步证明,使用Sidon清洗后的自动语音识别语料库训练的TTS模型,在零样本设置下可生成更高质量的合成语音。代码与模型已开源,以支持研究社区的可复现数据清洗工作。

原文摘要 · Abstract (English)

Large-scale text-to-speech (TTS) systems are limited by the scarcity of clean, multilingual recordings. We introduce Sidon, a fast, open-source speech restoration model that converts noisy in-the-wild speech into studio-quality speech and scales to dozens of languages. Sidon consists of two models: w2v-BERT 2.0 finetuned feature predictor to cleanse features from noisy speech and vocoder trained to synthesize restored speech from the cleansed features. Sidon achieves restoration performance comparable to Miipher: Google's internal speech restoration model with the aim of dataset cleansing for speech synthesis. Sidon is also computationally efficient, running up to 500 times faster than real time on a single GPU. We further show that training a TTS model using a Sidon-cleansed automatic speech recognition corpus improves the quality of synthetic speech in a zero-shot setting. Code and model are released to facilitate reproducible dataset cleansing for the research community.

语音修复多语言TTS高效计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。