arXiv:2412.14890eess.AScs.SD2024-12被引 2

通过可控数据集分析,发现语音增强模型更依赖声学属性而非语义属性。

Scale This, Not That: Investigating Key Dataset Attributes for Efficient Speech Enhancement Scaling

  • 用零样本语音合成生成可控数据集,分离并测试各属性影响。
  • 实验证明声学属性(如说话人、噪声)对性能提升远超语义属性。
  • 适合关注模型可扩展性与数据设计的研究者参考。

近期语音增强模型通过扩大模型规模和训练数据量取得了显著性能提升,但数据集变异性(如文本、语言、说话人、噪声)的影响仍缺乏深入研究。由于常用数据集中多种属性常相互纠缠,单独分析各属性影响极为困难。为此,我们提出一种生成-训练-评估框架,利用零样本文语转换系统,可控地生成训练数据并分别考察各属性变化对语音增强性能的影响。该方法可大规模合成数据,并精确控制单一属性。基于此框架,我们在多领域语料上分析了各类数据属性对判别式与生成式语音增强模型的缩放效应。大量实验表明,声学属性(如说话人、噪声)对当前语音增强模型的性能贡献远大于语义属性(如语言、文本),为未来研究提供了新视角。

原文摘要 · Abstract (English)

Recent speech enhancement models have shown impressive performance gains by scaling up model complexity and training data. However, the impact of dataset variability (e.g. text, language, speaker, and noise) has been underexplored. Analyzing each attribute individually is often challenging, as multiple attributes are usually entangled in commonly used datasets, posing a significant obstacle in understanding the distinct contributions of each attribute to the model's performance. To address this challenge, we propose a generation-training-evaluation framework that leverages zero-shot text-to-speech systems to investigate the impact of controlled attribute variations on speech enhancement performance. It enables us to synthesize training datasets in a scalable manner while carefully altering each attribute. Based on the proposed framework, we analyze the scaling effects of various dataset attributes on the performance of both discriminative and generative SE models. Extensive experiments on multi-domain corpora imply that acoustic attributes (e.g., speaker and noise) are much more important to current speech enhancement models than semantic attributes (e.g., language and text), offering new insights for future research.

语音增强数据属性可扩展性生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。