arXiv:2605.25512eess.AS2026-05被引 1

提出统一的复球面t分布混合模型,提升麦克风阵列语音分离性能。

cSTMM: A Unified Complex Spherical Student's $t$ Mixture Model for Directional Statistics in Mask-Based Blind Speech Separation

  • 用自由度和特征值约束统一多种复球面分布模型。
  • 在18种条件下平均信噪比提升0.25dB,最优自由度ν=1。
  • 适合需要高精度方向统计建模的语音分离场景。

基于方向统计的掩码盲语音分离(BSS)将多麦克风的归一化时频观测点聚类在复单位球面上,无需平面波或球面波假设。现有方法使用独立定义的角度混合模型,导致密度形状影响难以分离。本文提出复球面学生t混合模型(cSTMM),通过自由度ν和特征值约束连接复角中心高斯混合模型(cACGMM)、复宾汉姆混合模型(cBMM)和复沃森混合模型(cWMM)。我们推导了基于高浓度近似的隐变量-期望最大化(EM)框架,其中包含近似M步。在无噪声的LibriSpeech混音(经实测房间冲激响应混响)上,选定自由度ν*=1优于等价于cACGMM的ν=M,在全部18种测试条件下均表现更优,平均信干比改善(SDRi)达0.25 dB。该模型在ν=M时退化为cACGMM,大ν极限下趋近于cBMM/cWMM。

原文摘要 · Abstract (English)

Directional-statistics-based mask-based blind speech separation (BSS) clusters normalized time-frequency (TF) observations from $M$ microphones on the complex unit sphere, without relying on plane-wave or spherical-wave assumptions. Existing methods use separately defined angular mixture models, which makes the effect of density shape difficult to isolate. This paper proposes a complex spherical Student's $t$ mixture model (cSTMM) that connects the complex angular central Gaussian mixture model (cACGMM), complex Bingham mixture model (cBMM), and complex Watson mixture model (cWMM) through the degrees of freedom $ν$ and eigenvalue constraints. We derive a latent-scale expectation-maximization (EM) framework with an approximate M-step based on high-concentration approximation (HCA). On noise-free LibriSpeech mixtures reverberated using measured room impulse responses (RIRs), the development-selected value $ν^\ast=1$ outperformed the cACGMM-equivalent choice $ν=M$ in all 18 test conditions, yielding a mean signal-to-distortion ratio improvement (SDRi) gain of $0.25\,\mathrm{dB}$. The model reduces to the cACGMM at $ν=M$ and approaches the cBMM/cWMM in the large-$ν$ limits.

语音分离方向统计混合模型复球面

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。