arXiv:2410.11453eess.AScs.SD2024-10

融合空间与频谱信息提升多人语音定位跟踪精度

The importance of spatial and spectral information in multiple speaker tracking

  • 基于联合时空频信息的关联算法,提升说话人匹配准确率
  • 在LOCATA数据集上,联合信息使跟踪性能显著优于单一信息源
  • 适合需要高精度语音定位的智能会议系统、安防监控场景

使用麦克风阵列进行多说话人定位与跟踪在诸多应用中具有重要意义。多说话人跟踪的一大挑战是将方向估计正确关联到对应说话人。现有方法大多仅依赖空间或频谱信息,当其中任一信息不完整或缺失时性能下降。本文提出一种基于联合概率数据关联(JPDA)的方法,通过引入基于频谱信息估计的说话人时频掩码,参与关联概率计算,实现空间与频谱信息的联合利用。在LOCATA挑战数据集上的实验表明,采用联合时空频信息可显著提升跟踪性能。

原文摘要 · Abstract (English)

Multi-speaker localization and tracking using microphone array recording is of importance in a wide range of applications. One of the challenges with multi-speaker tracking is to associate direction estimates with the correct speaker. Most existing association approaches rely on spatial or spectral information alone, leading to performance degradation when one of these information channels is partially known or missing. This paper studies a joint probability data association (JPDA)-based method that facilitates association based on joint spatial-spectral information. This is achieved by integrating speaker time-frequency (TF) masks, estimated based on spectral information, in the association probabilities calculation. An experimental study that tested the proposed method on recordings from the LOCATA challenge demonstrates the enhanced performance obtained by using joint spatial-spectral information in the association.

语音跟踪麦克风阵列时频掩码多说话人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。