用语言描述分离音频,融合预训练模型提升精度
Hybrid-Sep: Language-queried audio source separation via pre-trained Model Fusion and Adversarial Diffusion Training
- 融合自监督音频模型与跨模态语义嵌入,实现精准音源分离
- 引入对抗一致性训练,在多个指标上超越现有最佳方法
- 适合需要高精度语言控制音频分离的研究者与开发者
语言查询音频分离(LASS)通过语义描述定位目标声音。现有方法在对齐复杂声学特征与语言上下文时面临挑战,且难以兼顾分离精度。当前研究多集中于文本描述增强与架构创新,而对预训练自监督学习(SSL)音频模型和对比语言-音频预训练(CLAP)框架的潜力挖掘不足。为此,我们提出HybridSep,一种两阶段LASS框架,融合基于SSL的声学表征与CLAP生成的语义嵌入。该框架引入对抗一致训练(ACT),将扩散过程作为辅助正则化损失,并结合对抗训练提升分离保真度。实验表明,HybridSep在多个指标上显著优于当前最优基线(如AudioSep、FlowSep),为LASS任务树立了新基准。
原文摘要 · Abstract (English)
Language-queried Audio Separation (LASS) employs linguistic queries to isolate target sounds based on semantic descriptions. However, existing methods face challenges in aligning complex auditory features with linguistic context while preserving separation precision. Current research efforts focus primarily on text description augmentation and architectural innovations, yet the potential of integrating pre-trained self-supervised learning (SSL) audio models and Contrastive Language-Audio Pretraining (CLAP) frameworks, capable of extracting cross-modal audio-text relationships, remains underexplored. To address this, we present HybridSep, a two-stage LASS framework that synergizes SSL-based acoustic representations with CLAP-derived semantic embeddings. Our framework introduces Adversarial Consistent Training (ACT), a novel optimization strategy that treats diffusion as an auxiliary regularization loss while integrating adversarial training to enhance separation fidelity. Experiments demonstrate that HybridSep achieves significant performance improvements over state-of-the-art baselines (e.g., AudioSep, FlowSep) across multiple metrics, establishing new benchmarks for LASS tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。