arXiv:2504.03998cs.SDeess.AS2025-04

建模语音频段相邻相关性,提升盲源分离效果

Determined blind source separation via modeling adjacent frequency band correlations in speech signals

  • 引入频段间相邻相关性约束,改进低秩矩阵分析
  • 在真实语音数据上实现更高信干比与失真比
  • 适合语音分离、声学信号处理研究者参考

多通道盲源分离(MBSS)在语音处理中被广泛研究。现有方法如独立低秩矩阵分析(ILRMA)和多通道非负矩阵分解(MNMF)利用源信号的低秩结构,但假设频段间相互独立;而独立向量分析(IVA)虽不依赖低秩模型,却基于均匀相关性假设。本文发现典型语音信号中相邻频段的相关性显著强于远距离频段。为此,提出基于加权Sinkhorn散度的ILRMA(wsILRMA),同时建模频段间依赖关系与联合概率分布。引入频段相关性约束后,相比现有方法在信噪比(SDR)和信干比(SIR)上均有提升。

原文摘要 · Abstract (English)

Multichannel blind source separation (MBSS), which focuses on separating signals of interest from mixed observations, has been extensively studied in acoustic and speech processing. Existing MBSS algorithms, such as independent low-rank matrix analysis (ILRMA) and multichannel nonnegative matrix factorization (MNMF), utilize the low-rank structure of source models but assume that frequency bins are independent. In contrast, independent vector analysis (IVA) does not rely on a low-rank source model but rather captures frequency dependencies based on a uniform correlation assumption. In this work, we demonstrate that dependencies between adjacent frequency bins are significantly stronger than those between bins that are farther apart in typical speech signals. To address this, we introduce a weighted Sinkhorn divergence-based ILRMA (wsILRMA) that simultaneously captures these inter-frequency dependencies and models joint probability distributions. Our approach incorporates an inter-frequency correlation constraint, leading to improved source separation performance compared to existing methods, as evidenced by higher Signal-to-Distortion Ratios (SDRs) and Source-to-Interference Ratios (SIRs).

盲源分离语音处理频段相关性信号分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。