arXiv:2601.22599cs.SDcs.HC2026-01中稿 · ICML被引 2

构建高纯度音效数据集,用更少数据实现更好声音分离效果。

A Semantically Consistent Dataset for Data-Efficient Query-Based Universal Sound Separation

  • 通过语义一致合成法剔除混音中事件共现,提升标签纯净度。
  • 新数据集Hive仅2.4千小时,性能超500倍更大的基准数据集模型。
  • 模型零样本泛化能力强,适合高效训练鲁棒听觉大模型。

基于查询的通用声音分离是智能听觉系统的核心,旨在从混合声中分离特定声源。尽管近期进展显著,现有方法在复杂声学场景下仍存在残留干扰。其主要瓶颈源于数据缺陷:真实世界数据集标签弱且事件严重共现,导致模型学习到背景噪声与目标类别间的虚假关联,而非稳健的声学特征。为此,我们提出自动化流水线,通过语义一致性合成协议,从真实数据中挖掘高纯净度单事件片段,消除事件共现。基于此,构建了名为Hive的高质量合成数据集,包含2.4千小时原始音频。实验表明,某些开源模型在仅使用Hive训练时,分离精度和感知质量已可媲美在约500倍更大数据集(~500小时)上训练的SOTA模型SAM-Audio;同时在分布外评估基准上展现出优异的零样本泛化能力。结果表明,优先保障监督信号纯净度可显著提升数据效率,为以更低计算成本训练鲁棒听觉基础模型提供新范式。代码与数据集详见https://cslikai.cn/Hive。

原文摘要 · Abstract (English)

Query-based universal sound separation is fundamental to intelligent auditory systems, aiming to isolate specific sources from mixtures. Despite recent advances, existing methods continue to suffer from residual interference in complex acoustic scenes. This performance limitation stems largely from a data bottleneck: in-the-wild datasets contain weak labels and severe co-occurrence of events. These flaws induce models to learn spurious correlations between background noise and target categories instead of robust acoustic features. To address this, we propose an automated pipeline that eliminates co-occurrence of events by mining high-purity single-event segments from in-the-wild datasets via a semantically consistent synthesis protocol. Utilizing this pipeline, we constructed Hive, a high-quality synthetic dataset comprising 2.4k hours of raw audio. Experimental results demonstrate that, compared with the state-of-the-art model SAM-Audio which was trained on a huge dataset $\sim$500 times larger than Hive, certain open-source models trained on Hive achieve competitive separation accuracy and perceptual quality. Moreover, these models exhibited remarkable zero-shot generalization on out-of-distribution evaluation benchmarks. These findings highlight that prioritizing purity of supervised signals enables significant data efficiency, offering a new paradigm for training robust auditory foundation models with reduced computational costs. Code and dataset are available at https://cslikai.cn/Hive.

声音分离数据集零样本高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。