arXiv:2510.03111eess.AS2025-10

对比24种预处理方案,找出最适合野生语音合成的数据集构建方法

Evaluation of preprocessing pipelines in the creation of in-the-wild TTS datasets

  • 用客观指标对比24种降噪与筛选组合,评估效果
  • 发现宽松筛选搭配降噪可平衡数据量与语音质量
  • 无需训练模型即可选最优方案,适合资源有限场景

本文提出一种可复现、基于度量的评估方法,用于检验野生环境语音合成语料库的预处理流程。我们采用低成本自研流程处理首个阿根廷西班牙语野生语料库,并对比了24种结合不同降噪与质量过滤策略的配置方案。评估采用互补的客观指标(PESQ、SI-SDR、SNR)、声学描述符(T30、C50)以及语音保真度指标(F0-STD、MCD)。结果揭示了数据集规模、信号质量与语音保真度之间的权衡关系;在测试环境中,采用宽松过滤的降噪方案表现最佳。该方法可在不训练语音合成模型的前提下选择最优配置,显著加速并降低低资源环境下预处理开发的成本。

原文摘要 · Abstract (English)

This work introduces a reproducible, metric-driven methodology to evaluate preprocessing pipelines for in-the-wild TTS corpora generation. We apply a custom low-cost pipeline to the first in-the-wild Argentine Spanish collection and compare 24 pipeline configurations combining different denoising and quality filtering variants. Evaluation relies on complementary objective measures (PESQ, SI-SDR, SNR), acoustic descriptors (T30, C50), and speech-preservation metrics (F0-STD, MCD). Results expose trade-offs between dataset size, signal quality, and voice preservation; where denoising variants with permissive filtering provide the best overall compromise for our testbed. The proposed methodology allows selecting pipeline configurations without training TTS models for each subset, accelerating and reducing the cost of preprocessing development for low-resource settings.

语音合成数据预处理低资源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。