arXiv:2602.08671eess.AS2026-02中稿 · IEEE TASLP被引 1

用序列建模实现自适应频谱压缩,提升音源分离效率与性能

Input-Adaptive Spectral Feature Compression by Sequence Modeling for Source Separation

  • 用单个序列模型替代固定分带编码,实现输入自适应压缩
  • 在音乐和影院音频分离任务中,性能全面优于传统分带模块
  • 支持不同模型规模与压缩率,适合高采样率音频处理场景

时频域双路径模型在音源分离中表现优异,但计算开销随频带数量增长。在高采样率任务(如音乐源分离MSS、影院音频源分离CASS)中,常采用分带(BS)模块压缩频谱信息,通过为每个预定义子带独立编码实现有效压缩,并引入强调低频的归纳偏置。然而,该模块存在两大固有缺陷:(i) 非输入自适应,无法利用输入依赖信息;(ii) 参数量大,每个子带需独立模块。为此,本文提出频谱特征压缩(SFC),使用单一序列建模模块实现输入自适应与参数高效压缩。研究了基于交叉注意力与Mamba的两种SFC变体,并引入受BS启发的归纳偏置,使其适用于频谱信息压缩。在MSS与CASS任务上的实验表明,SFC在不同分离器规模与压缩比下均持续优于BS模块。分析还显示,SFC能自适应捕捉输入中的频率模式。

原文摘要 · Abstract (English)

Time-frequency domain dual-path models have demonstrated strong performance and are widely used in source separation. Because their computational cost grows with the number of frequency bins, these models often use the band-split (BS) module in high-sampling-rate tasks such as music source separation (MSS) and cinematic audio source separation (CASS). The BS encoder compresses frequency information by encoding features for each predefined subband. It achieves effective compression by introducing an inductive bias that places greater emphasis on low-frequency parts. Despite its success, the BS module has two inherent limitations: (i) it is not input-adaptive, preventing the use of input-dependent information, and (ii) the parameter count is large, since each subband requires a dedicated module. To address these issues, we propose Spectral Feature Compression (SFC). SFC compresses the input using a single sequence modeling module, making it both input-adaptive and parameter-efficient. We investigate two variants of SFC, one based on cross-attention and the other on Mamba, and introduce inductive biases inspired by the BS module to make them suitable for frequency information compression. Experiments on MSS and CASS tasks demonstrate that the SFC module consistently outperforms the BS module across different separator sizes and compression ratios. We also provide an analysis showing that SFC adaptively captures frequency patterns from the input.

音源分离频谱压缩序列建模Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。