arXiv:2409.15885cs.SDcs.LG2024-09被引 1

用置信度识别高错误区域,提升语音分离模型校准与标注效率。

On the calibration of powerset speaker diarization models

  • 基于置信度筛选低置信区域进行训练和验证。
  • 在多个数据集上,低置信区训练使模型更校准,标注效率更高。
  • 高置信度可有效预测错误集中区域,适合优化标注策略。

端到端神经语音分离模型通常采用多标签分类方法。我们此前提出一种幂集多类公式,在多个数据集上超越了现有最优结果。本文研究该幂集模型的校准性,涵盖域内与域外情况,并分析低置信区域的数据特征。通过测试模型置信度的实际可靠性,我们利用预训练模型的置信度从无标注数据中选择性构建训练与验证子集,对比随机选取。结果表明,顶标签置信度可可靠预测高误差区域;在低置信区域训练能获得更校准的模型;在低置信区域验证比随机区域更具标注效率。

原文摘要 · Abstract (English)

End-to-end neural diarization models have usually relied on a multilabel-classification formulation of the speaker diarization problem. Recently, we proposed a powerset multiclass formulation that has beaten the state-of-the-art on multiple datasets. In this paper, we propose to study the calibration of a powerset speaker diarization model, and explore some of its uses. We study the calibration in-domain, as well as out-of-domain, and explore the data in low-confidence regions. The reliability of model confidence is then tested in practice: we use the confidence of the pretrained model to selectively create training and validation subsets out of unannotated data, and compare this to random selection. We find that top-label confidence can be used to reliably predict high-error regions. Moreover, training on low-confidence regions provides a better calibrated model, and validating on low-confidence regions can be more annotation-efficient than random regions.

语音分离模型校准标注效率置信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。