arXiv:2511.22696cs.SDcs.AI2025-11被引 1

提出概率级融合与校准框架,提升语音分离模型的可靠性与性能。

Probabilistic Fusion and Calibration of Neural Speaker Diarization Models

  • 基于连续概率输出实现更优的模型融合与校准
  • 在CallHome数据集上实现最高19%相对DER降低
  • 适合需要可信置信度的下游语音应用

端到端神经语音分离(EEND)系统生成帧级说话人活动概率估计,但因评估主要关注语音分离错误率(DER),其置信度得分的可靠性与校准常被忽视。当前唯一成熟的方法DOVER-Lap在片段层面进行硬决策。本文首次提出在概率层面进行融合与校准的完整框架,研究了多标签与幂集表示两种输出形式对校准与融合效果的影响。在CallHome双说话人基准上,实验表明恰当校准可显著提升单个模型性能(最多19%相对DER下降),在某些情况下可弥补领域适应缺失。结果表明:幂集空间联合校准优于逐说话人独立校准;融合显著优于单个模型;先融合后校准的顺序优于先校准再融合或未校准融合,且仅需校准单一组合模型。最优配置在DER上超越DOVER-Lap,并提供可靠置信度,适用于下游任务。本工作为EEND系统的概率级融合提供最佳实践。

原文摘要 · Abstract (English)

End-to-End Neural Diarization (EEND) systems produce frame-level probabilistic speaker activity estimates, yet since evaluation focuses primarily on Diarization Error Rate (DER), the reliability and calibration of these confidence scores have been largely neglected. When fusing multiple diarization systems, DOVER-Lap remains the only established approach, operating at the segment level with hard decisions. We propose working with continuous probability outputs, which enables more sophisticated fusion and calibration techniques that can leverage model uncertainty and complementary strengths across different architectures. This paper presents the first comprehensive framework for calibrating and fusing EEND models at the probability level. We investigate two output formulations (multilabel and powerset representations) and their impact on calibration and fusion effectiveness. Through extensive experiments on the CallHome two-speaker benchmark, we demonstrate that proper calibration provides substantial improvements even for individual models (up to 19% relative DER reduction), in some cases mitigating the absence of domain adaptation. We reveal that joint calibration in powerset space consistently outperforms independent per-speaker calibration, that fusion substantially improves over individual models, and that the Fuse-then-Calibrate ordering generally outperforms both calibrating before fusion and uncalibrated fusion while requiring calibration of only a single combined model. Our best configuration outperforms DOVER-Lap in terms of DER while providing reliable confidence estimates essential for downstream applications. This work proposes best practices for probability-level fusion of EEND systems and demonstrates the advantages of leveraging soft outputs over hard decisions.

语音分离概率融合模型校准置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。