arXiv:2608.28661cs.SD2026-08中稿 · Interspeech 2026

用贝塔先验提升远场语音分离与说话人辨认的鲁棒性

Neural Multichannel Distant Speaker Diarization and Source Separation with Beta Speaker Activity Prior

论文配图:Neural Multichannel Distant Speaker Diarization and Source Separation with Beta Speaker Activity Prior
图 1 · 摘自论文原文
  • 引入贝塔分布先验约束说话人活跃度,替代传统交叉熵损失
  • 在AMI数据集上纠错率降低至少3%(相对16%),杰卡德误差降4%(相对20%)
  • 适合研究多通道语音处理、需提升模型稳健性的工程师

远场说话人辨认因声学环境恶劣、说话人数变化和语音重叠而极具挑战。尽管数据驱动方法表现优异,模型驱动方法通过利用多通道录音的空间信息仍具吸引力。本文提出一种基于贝叶斯框架的模型驱动方法——神经FCASA,通过在说话人活跃度上引入贝塔先验,构建可优化的变分下界目标函数,以连续说话人活跃度得分作为正则化替代原交叉熵损失来训练辨认模型。实验表明,在AMI数据集上,该方法在说话人辨认错误率(DER)上至少提升3%(相对16%),杰卡德错误率(JER)至少提升4%(相对20%),显著优于基线。

原文摘要 · Abstract (English)

Distant speaker diarization remains challenging due to adverse acoustic conditions, varying numbers of speakers and overlapping speech. While data-driven approaches have shown strong performance, model-driven methods offer a compelling alternative by leveraging spatial information from multichannel recordings. This paper is motivated to propose a Bayesian diarization model for a model-driven method called neural FCASA to enhance its robustness. Specifically, we propose a beta prior over speaker activity and hence a variational lower bound objective that can be seen as a regularized continuous speaker activity score in place of the original cross-entropy loss to train the diarization model. Our experiments show significant improvements in terms of Diarization Error Rate by at least 3% (16% relatively) and Jaccard Error Rate by at least 4% (20% relatively) on the AMI dataset compared to the baseline.

说话人辨认多通道音频贝叶斯建模语音分离

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。