提出松耦合空间谱模型,提升动态环境下会议语音分离与说话人识别效果。
Loose coupling of spectral and spatial models for multi-channel diarization and enhancement of meetings in dynamic environments
- 用概率建模声源位置与说话人对应关系,实现空间与谱信息松耦合
- 在LibriCSS数据集上,动态位置变化下性能优于紧耦合系统
- 适合移动说话人场景的会议语音处理,尤其适用于麦克风阵列应用
麦克风阵列捕获声音可同时利用空间和频谱信息进行说话人分离与信号增强,这两项任务对会议转录至关重要。然而,当说话人移动时,空间位置与说话人之间不存在一一对应关系。为此,本文提出一种新的联合空间-谱混合模型,其两个子模型通过概率化建模说话人与位置索引的关系实现松耦合。该方法既能联合利用空间与谱信息,又允许同一说话人从不同位置发声。在模拟说话人位置变化的LibriCSS数据集上的实验表明,该方法相比紧耦合系统有显著性能提升。
原文摘要 · Abstract (English)
Sound capture by microphone arrays opens the possibility to exploit spatial, in addition to spectral, information for diarization and signal enhancement, two important tasks in meeting transcription. However, there is no one-to-one mapping of positions in space to speakers if speakers move. Here, we address this by proposing a novel joint spatial and spectral mixture model, whose two submodels are loosely coupled by modeling the relationship between speaker and position index probabilistically. Thus, spatial and spectral information can be jointly exploited, while at the same time allowing for speakers speaking from different positions. Experiments on the LibriCSS data set with simulated speaker position changes show great improvements over tightly coupled subsystems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。