统一模型同时解决声音分离与目标提取,自动判断声源数量并响应多种用户提示。
USE: A Unified Model for Universal Sound Separation and Extraction
- 设计双模块架构:自适应声源数推断+多模态线索融合
- 声音分离提升1.4 dB SDR,目标提取准确率达86%
- 支持全自动分离或提示驱动提取,适用复杂听觉场景
声音分离(SS)和目标声音提取(TSE)是应对复杂声学场景的基础技术。现有SS方法难以确定未知声源数量,TSE则需精确提示才能达到最佳性能。本文提出一个统一框架,协同结合SS与TSE以克服各自局限。架构包含两个互补组件:1)编码器-解码器吸引子(EDA)网络,可自动推断声源数量及对应声学线索用于SS;2)多模态融合网络,精准解析用户提供的多样化线索(声学、语义或视觉)用于TSE。通过跨任务一致性约束联合训练,建立统一潜在空间以连接两种范式。推理时系统可自适应切换至全自主SS模式或提示驱动的TSE模式。实验表明,在两项任务中均表现优异,相比基线在SS上提升1.4 dB SDR,TSE准确率达86%。
原文摘要 · Abstract (English)
Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require precisely specified clues to achieve optimal performance. This paper proposes a unified framework that synergistically combines SS and TSE to overcome their individual limitations. Our architecture employs two complementary components: 1) An Encoder-Decoder Attractor (EDA) network that automatically infers both the source count and corresponding acoustic clues for SS, and 2) A multi-modal fusion network that precisely interprets diverse user-provided clues (acoustic, semantic, or visual) for TSE. Through joint training with cross-task consistency constraints, we establish a unified latent space that bridges both paradigms. During inference, the system adaptively operates in either fully autonomous SS mode or clue-driven TSE mode. Experiments demonstrate remarkable performance in both tasks, with notable improvements of 1.4 dB SDR improvement in SS compared to baseline and 86\% TSE accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。