arXiv:2604.12480cs.SDcs.AI2026-04被引 7

用β散度非负分解提升混响环境下的音频分离效果

Audio Source Separation in Reverberant Environments using $β$-divergence based Nonnegative Factorization

论文配图:Audio Source Separation in Reverberant Environments using $β$-divergence based Nonnegative Factorization
图 1 · 摘自论文原文
  • 基于先验谱方差的非负分解,通过β散度优化参数估计
  • 在多种混响条件下,分离性能优于现有方法
  • 适合需要高保真音频分离的语音/音乐处理场景

在基于高斯模型的多通道音频源分离中,观测混合信号的似然由源信号频谱方差和对应的空间协方差矩阵参数化。这些参数通过期望最大化算法最大化似然来估计,并用于多通道维纳滤波实现信号分离。本文提出通过结合源方差先验信息,采用基于β散度的非负分解来估计这些参数。谱基矩阵可作为先验信息定义,既可直接提取,也可通过预先训练的冗余库间接获得。通过独立的非负张量分解步骤,提出了两种算法以提取或检测最能表征观测混合信号功率谱的基矩阵。分解通过最小化β散度并使用乘法更新规则实现,因子分解的稀疏性可通过调节β值控制。实验表明,稀疏性而非β值本身对提升分离性能更为关键。所提方法在多种混合条件下进行了评估,结果优于其他可比算法。

原文摘要 · Abstract (English)

In Gaussian model-based multichannel audio source separation, the likelihood of observed mixtures of source signals is parametrized by source spectral variances and by associated spatial covariance matrices. These parameters are estimated by maximizing the likelihood through an Expectation-Maximization algorithm and used to separate the signals by means of multichannel Wiener filtering. We propose to estimate these parameters by applying nonnegative factorization based on prior information on source variances. In the nonnegative factorization, spectral basis matrices can be defined as the prior information. The matrices can be either extracted or indirectly made available through a redundant library that is trained in advance. In a separate step, applying nonnegative tensor factorization, two algorithms are proposed in order to either extract or detect the basis matrices that best represent the power spectra of the source signals in the observed mixtures. The factorization is achieved by minimizing the $β$-divergence through multiplicative update rules. The sparsity of factorization can be controlled by tuning the value of $β$. Experiments show that sparsity, rather than the value assigned to $β$ in the training, is crucial in order to increase the separation performance. The proposed method was evaluated in several mixing conditions. It provides better separation quality with respect to other comparable algorithms.

音频分离非负分解混响环境β散度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。