用说话人嵌入修复间歇移动说话人的追踪身份错乱问题
Speaker Embeddings to Improve Tracking of Intermittent and Moving Speakers
- 后处理阶段利用说话人嵌入重分配追踪轨迹身份
- 在说话人位置变化时仍保持90%以上的身份识别准确率
- 适合语音会议、智能音箱等动态场景的说话人追踪
说话人追踪通常依赖空间观测来维持身份连贯性,但在说话人间歇且移动的场景下(如静默时改变位置),会导致空间轨迹不连续。本文提出一种简单方案:在初始追踪后,利用多通道音频信号通过波束成形增强目标方向语音,提取说话人嵌入,并基于注册池进行身份重分配。实验在包含静默期位置变化的语料库上验证,该方法显著提升神经与传统追踪系统的身份分配性能。研究还分析了波束成形和嵌入提取时长的影响,结果表明短时(1.5秒)波束成形即可实现稳定效果。
原文摘要 · Abstract (English)
Speaker tracking methods often rely on spatial observations to assign coherent track identities over time. This raises limits in scenarios with intermittent and moving speakers, i.e., speakers that may change position when they are inactive, thus leading to discontinuous spatial trajectories. This paper proposes to investigate the use of speaker embeddings, in a simple solution to this issue. We propose to perform identity reassignment post-tracking, using speaker embeddings. We leverage trajectory-related information provided by an initial tracking step and multichannel audio signal. Beamforming is used to enhance the signal towards the speakers' positions in order to compute speaker embeddings. These are then used to assign new track identities based on an enrollment pool. We evaluate the performance of the proposed speaker embedding-based identity reassignment method on a dataset where speakers change position during inactivity periods. Results show that it consistently improves the identity assignment performance of neural and standard tracking systems. In particular, we study the impact of beamforming and input duration for embedding extraction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。