arXiv:2601.12254cs.SDeess.AS2026-01中稿 · ICASSP 2026

用生成令牌概率检测语音增强中的幻觉错误,提升数据集质量

Confidence-based Filtering for Speech Dataset Curation with Generative Speech Enhancement Using Discrete Tokens

  • 用生成令牌的对数概率作置信度,非侵入式检测幻觉错误
  • 在多个指标上表现优于传统方法,能发现被忽略的错误
  • 适合用于真实场景语音合成数据集的清洗与优化

生成式语音增强(GSE)模型能从噪声输入中生成高质量清洁语音,适用于将嘈杂的文本到语音(TTS)数据集转换为高质量数据集。然而,这类模型容易产生幻觉错误,如音素遗漏和说话人不一致,而基于非侵入式语音质量指标的传统错误过滤方法往往无法识别这些错误。为此,我们提出一种针对基于离散令牌的GSE模型的非侵入式幻觉错误过滤方法。该方法利用生成令牌的对数概率作为置信度分数,以检测潜在错误。实验表明,置信度分数与一系列侵入式语音增强指标具有强相关性,且能有效识别传统方法遗漏的幻觉错误。此外,我们展示了该方法的实际应用价值:使用该置信度过滤方法清洗真实场景下的TTS数据集后,后续训练的TTS模型性能显著提升。

原文摘要 · Abstract (English)

Generative speech enhancement (GSE) models show great promise in producing high-quality clean speech from noisy inputs, enabling applications such as curating noisy text-to-speech (TTS) datasets into high-quality ones. However, GSE models are prone to hallucination errors, such as phoneme omissions and speaker inconsistency, which conventional error filtering based on non-intrusive speech quality metrics often fails to detect. To address this issue, we propose a non-intrusive method for filtering hallucination errors from discrete token-based GSE models. Our method leverages the log-probabilities of generated tokens as confidence scores to detect potential errors. Experimental results show that the confidence scores strongly correlate with a suite of intrusive SE metrics, and that our method effectively identifies hallucination errors missed by conventional filtering methods. Furthermore, we demonstrate the practical utility of our method: curating an in-the-wild TTS dataset with our confidence-based filtering improves the performance of subsequently trained TTS models.

语音增强数据清洗生成模型置信度过滤

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。