arXiv:2504.17782cs.SDcs.LG2025-04被引 1

用真实混音数据训练,让声音分离模型更适应现实环境。

Unleashing the Power of Natural Audio Featuring Multiple Sound Sources

  • 通过数据引擎拆解真实混音,生成独立音轨用于训练。
  • 在多个任务上达到当前最优性能,超越传统方法。
  • 适合需要真实场景声音分离的应用开发者。

通用声音分离旨在从混合音频中提取对应不同事件的纯净音频轨道,对人工听觉感知至关重要。然而,现有方法严重依赖人工混合音频训练,限制了其在真实环境中自然混音上的泛化能力。为此,我们提出ClearSep框架,利用数据引擎将复杂自然混音分解为多个独立音轨,从而实现真实场景下的有效声音分离。引入基于重混的评估指标,量化分离质量,并将其作为阈值,迭代应用数据引擎与模型训练,持续优化分离性能。此外,设计一系列针对分离后独立音轨的训练策略,以最大化利用这些数据。大量实验表明,ClearSep在多个声音分离任务上均达到最先进水平,展现了其在自然音频场景下推进声音分离的潜力。更多示例和详细结果请访问我们的演示页面:https://clearsep.github.io。

原文摘要 · Abstract (English)

Universal sound separation aims to extract clean audio tracks corresponding to distinct events from mixed audio, which is critical for artificial auditory perception. However, current methods heavily rely on artificially mixed audio for training, which limits their ability to generalize to naturally mixed audio collected in real-world environments. To overcome this limitation, we propose ClearSep, an innovative framework that employs a data engine to decompose complex naturally mixed audio into multiple independent tracks, thereby allowing effective sound separation in real-world scenarios. We introduce two remix-based evaluation metrics to quantitatively assess separation quality and use these metrics as thresholds to iteratively apply the data engine alongside model training, progressively optimizing separation performance. In addition, we propose a series of training strategies tailored to these separated independent tracks to make the best use of them. Extensive experiments demonstrate that ClearSep achieves state-of-the-art performance across multiple sound separation tasks, highlighting its potential for advancing sound separation in natural audio scenarios. For more examples and detailed results, please visit our demo page at https://clearsep.github.io.

声音分离真实音频数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。