arXiv:2506.23859eess.AScs.SD2025-06中稿 · ASRU2025

精炼数据比堆量更有效,700小时优质数据胜过2500小时杂乱数据。

Less is More: Data Curation Matters in Scaling Speech Enhancement

  • 精选700小时高质量语音数据,替代全量2500小时数据集。
  • 模型在700小时精选数据上表现优于2500小时原始数据集。
  • 适合关注语音增强数据质量与高效训练的研究者。

当前主流语音增强系统依赖数据驱动的神经网络模型。传统观点认为数据量越大性能越好,这一现象在多个领域得到验证。然而,近期研究表明语音增强任务中数据规模扩大存在边际收益递减。本文聚焦大规模数据集中‘干净’标签存在的普遍质量问题,重新审视该现象并证实:在大规模训练集中,优先选择高质量数据比单纯扩充数据量更为关键。实验表明,使用精心筛选的700小时数据训练的模型,性能超越在2500小时完整数据集上训练的模型。这一结果凸显了数据清洗在语音增强系统可扩展性中的核心作用。

原文摘要 · Abstract (English)

The vast majority of modern speech enhancement systems rely on data-driven neural network models. Conventionally, larger datasets are presumed to yield superior model performance, an observation empirically validated across numerous tasks in other domains. However, recent studies reveal diminishing returns when scaling speech enhancement data. We focus on a critical factor: prevalent quality issues in ``clean'' training labels within large-scale datasets. This work re-examines this phenomenon and demonstrates that, within large-scale training sets, prioritizing high-quality training data is more important than merely expanding the data volume. Experimental findings suggest that models trained on a carefully curated subset of 700 hours can outperform models trained on the 2,500-hour full dataset. This outcome highlights the crucial role of data curation in scaling speech enhancement systems effectively.

语音增强数据清洗模型效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。