融合声源定位与说话人嵌入,实现无需麦克风配置先验的会议语音分段聚类。
Spatio-spectral diarization of meetings by combining TDOA-based segmentation and speaker embedding-based clustering
- 先用时差定位分割语音段,再用嵌入向量聚类区分说话人。
- 在紧凑阵列和分布式麦克风下均显著优于单通道方法。
- 能准确追踪移动说话人,适合真实会议场景应用。
我们提出一种基于时空联合建模的会议语音分角色识别流程,包含基于时差定位(TDOA)的语音段分割与基于说话人嵌入的聚类。该系统无需多通道训练数据,也不依赖麦克风数量或布局的先验知识,适用于紧凑麦克风阵列和分布式麦克风部署,仅需微调即可适配。由于在分割阶段对重叠语音具有更强处理能力,该方案在紧凑阵列与分布式麦克风两种场景下均显著优于单通道pyannote方法。此外,与完全依赖空间信息的方法不同,本系统可在说话人位置变化时仍正确追踪其身份。
原文摘要 · Abstract (English)
We propose a spatio-spectral, combined model-based and data-driven diarization pipeline consisting of TDOA-based segmentation followed by embedding-based clustering. The proposed system requires neither access to multi-channel training data nor prior knowledge about the number or placement of microphones. It works for both a compact microphone array and distributed microphones, with minor adjustments. Due to its superior handling of overlapping speech during segmentation, the proposed pipeline significantly outperforms the single-channel pyannote approach, both in a scenario with a compact microphone array and in a setup with distributed microphones. Additionally, we show that, unlike fully spatial diarization pipelines, the proposed system can correctly track speakers when they change positions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。