用最优传输统一解决声音分离与到达时间差估计问题
Joint Spectrogram Separation and TDOA Estimation using Optimal Transport
- 基于最优传输理论构建联合分离与延迟估计框架
- 在多种噪声下对语音信号实现高精度分离与延迟估计
- 适合语音增强、多通道定位等实际场景应用
在语音增强和电信等领域,分离重叠声源是常见挑战,准确区分重叠声音有助于降低干扰并提升信号质量。在多通道系统中,正确校准与同步对精确分离和定位声源至关重要。本文提出一种盲源分离与时间差到达(TDOA)估计方法,该方法在时频域中同时实现信号混合的分离与接收器间相对延迟的估计。通过利用最优传输(OT)问题的结构,将分离与延迟估计整合为统一框架,并采用块坐标下降算法进行优化。我们在不同噪声条件下分析了基于OT的估计算法性能,并与传统TDOA和源分离方法进行了比较。数值仿真结果表明,该方法在多种噪声环境下对真实语音信号的TDOA估计和源分离任务均表现出显著的高精度。
原文摘要 · Abstract (English)
Separating sources is a common challenge in applications such as speech enhancement and telecommunications, where distinguishing between overlapping sounds helps reduce interference and improve signal quality. Additionally, in multichannel systems, correct calibration and synchronization are essential to separate and locate source signals accurately. This work introduces a method for blind source separation and estimation of the Time Difference of Arrival (TDOA) of signals in the time-frequency domain. Our proposed method effectively separates signal mixtures into their original source spectrograms while simultaneously estimating the relative delays between receivers, using Optimal Transport (OT) theory. By exploiting the structure of the OT problem, we combine the separation and delay estimation processes into a unified framework, optimizing the system through a block coordinate descent algorithm. We analyze the performance of the OT-based estimator under various noise conditions and compare it with conventional TDOA and source separation methods. Numerical simulation results demonstrate that our proposed approach can achieve a significant level of accuracy across diverse noise scenarios for physical speech signals in both TDOA and source separation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。