arXiv:2606.11795eess.AScs.SD2026-06中稿 · Interspeech 2026

用因果模型生成紧致语音段,提升说话人辨识精度

Tight Boundary Prediction in Speaker Diarization Using Causal-Anticausal Consistency

  • 用因果与反因果模型生成更紧的伪标签
  • 迭代优化使预测紧致度达理想训练的70%
  • 适合需要精确语音分段的下游任务

多说话人对话语音识别数据常用于训练说话人辨识模型。由于这类数据强调语义连续性,语音片段中包含停顿和边界余量,导致标注松散。基于此类数据训练的模型会内化这种松散特性,尽管在某些下游应用中更紧的语音区间更优。本文提出新任务:在使用松散标签的情况下实现紧致预测。方法通过因果与反因果模型生成紧致伪标签,二者均无法学习松散行为。进一步提出协同训练机制,迭代收紧标签并更新双模型以实现渐进优化。实验表明,该方法恢复了理想紧标签训练约70%的紧致效果,并提升了下游性能。

原文摘要 · Abstract (English)

Multi-talker conversational automatic speech recognition data are often used to train speaker diarization models. Because such data prioritize semantic continuity, pauses and boundary margins are included within speech segments, resulting in loose annotations. Models trained on such data tend to internalize mechanisms that reproduce this looseness, although tight speech intervals are sometimes preferable for downstream applications. In this paper, we address the novel task of enabling models to produce tight predictions using loose labels. Our method generates tighter pseudo labels using causal and anticausal models, which are inherently incapable of learning loosening behavior. We further propose a co-training scheme that iteratively tightens labels and updates both models for more progressive refinement. Experimental results show that the proposed method recovers about 70 % of the tightening effect achieved by ideal tight-label training and improves downstream performance.

说话人辨识语音分割自训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。