arXiv:2409.12413eess.AScs.SD2024-09被引 14

DeFT-Mamba统一解决多通道声音分离与多音符音频分类问题。

DeFT-Mamba: Universal Multichannel Sound Separation and Polyphonic Audio Classification

  • 结合DeFTAN与Mamba捕捉局部与全局时频关系
  • 在复杂多音场景下性能显著超越现有方法
  • 适合需要高精度声音分离的智能音频系统

本文提出一种通用多通道声音分离与多音符音频分类框架DeFT-Mamba,解决混合信号中分离与分类单个声源的难题。该框架采用密集时频注意力网络(DeFTAN)与Mamba结构,通过门控卷积模块捕捉局部时频关联,通过位置感知混合Mamba建模全局时频依赖。DeFT-Mamba在复杂同类多音场景下性能大幅超越现有分离与分类模型。同时引入基于分类的源数估计方法,优于传统阈值法。还提出分离优化调优策略进一步提升效果。框架在自建的多通道通用声音分离数据集上训练与测试,该数据集模拟真实环境中的移动声源及多音事件的起始与终止变化。

原文摘要 · Abstract (English)

This paper presents a framework for universal sound separation and polyphonic audio classification, addressing the challenges of separating and classifying individual sound sources in a multichannel mixture. The proposed framework, DeFT-Mamba, utilizes the dense frequency-time attentive network (DeFTAN) combined with Mamba to extract sound objects, capturing the local time-frequency relations through gated convolution block and the global time-frequency relations through position-wise Hybrid Mamba. DeFT-Mamba surpasses existing separation and classification networks by a large margin, particularly in complex scenarios involving in-class polyphony. Additionally, a classification-based source counting method is introduced to identify the presence of multiple sources, outperforming conventional threshold-based approaches. Separation refinement tuning is also proposed to improve performance further. The proposed framework is trained and tested on a multichannel universal sound separation dataset developed in this work, designed to mimic realistic environments with moving sources and varying onsets and offsets of polyphonic events.

声音分离多音音频Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。