arXiv:2603.02641cs.SD2026-03被引 4

提出新方法提升语音增强效果,兼顾清晰度与自然度。

Rethinking Training Targets, Architectures and Data Quality for Universal Speech Enhancement

  • 用时移无混响语音作训练目标,优于传统反射语音
  • 两阶段框架在感知质量下最小化失真,性能达最新水平
  • 数据质量比数量更重要,适合语音合成数据优化

通用语音增强(USE)旨在多种失真条件下恢复语音质量并保持信号保真度。尽管近期取得进展,但训练目标选择、失真-感知权衡及数据筛选等关键问题仍未解决。本文系统性地解决这三个被忽视的问题:首先,重新审视使用早期反射语音作为去混响目标的做法,发现其会降低感知质量和下游语音识别性能;改用时移的无混响干净语音作为学习目标可显著提升效果。其次,基于失真-感知权衡理论,提出简单高效的两阶段框架,在给定感知质量水平下实现最小失真。第三,分析了训练数据规模与质量之间的权衡,揭示在大规模未筛选语料上训练存在性能上限,因模型难以去除细微伪影。所提方法在URGENT 2025非盲测试集上达到当前最佳表现,并展现出强语言无关泛化能力,适用于提升语音合成训练数据质量。模型权重已公开于https://huggingface.co/nvidia/RE-USE。

原文摘要 · Abstract (English)

Universal Speech Enhancement (USE) aims to restore speech quality under diverse degradation conditions while preserving signal fidelity. Despite recent progress, key challenges in training target selection, the distortion--perception tradeoff, and data curation remain unresolved. In this work, we systematically address these three overlooked problems. First, we revisit the conventional practice of using early-reflected speech as the dereverberation target and show that it can degrade perceptual quality and downstream ASR performance. We instead demonstrate that time-shifted anechoic clean speech provides a superior learning target. Second, guided by the distortion--perception tradeoff theory, we propose a simple two-stage framework that achieves minimal distortion under a given level of perceptual quality. Third, we analyze the trade-off between training data scale and quality for USE, revealing that training on large uncurated corpora imposes a performance ceiling, as models struggle to remove subtle artifacts. Our method achieves state-of-the-art performance on the URGENT 2025 non-blind test set and exhibits strong language-agnostic generalization, making it effective for improving TTS training data. Model weights are available for download at: https://huggingface.co/nvidia/RE-USE.

语音增强语音质量数据质量语音合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。