arXiv:2512.15229cs.LGcs.SD2025-12

提出高效在线说话人分离系统,无需调参且计算成本低。

O-EENC-SD: Efficient Online End-to-End Neural Clustering for Speaker Diarization

  • 基于RNN的拼接机制实现在线预测,无需依赖聚类超参数。
  • 在CallHome数据集上达到与顶尖方法相当的错误率(DER)。
  • 适合实时语音分析场景,尤其擅长独立分块处理,效率极高。

我们提出O-EENC-SD:一种基于EEND-EDA的端到端在线说话人分离系统,采用新型基于RNN的拼接机制实现在线预测。特别地,我们设计了一种新的中心点优化解码器,并通过严格的消融实验验证其有效性。相比现有方法,该系统具有显著优势:相较于无监督聚类方法,无需调节超参数;相较于当前端到端在线方法,计算开销更低。我们在两说话人电话对话语料库上测试,结果表明O-EENC-SD在呼叫中心(CallHome)数据集上性能可媲美最先进水平。实验显示,即便在无重叠的独立分块输入下,该系统仍能实现错误率(DER)与复杂度之间的良好平衡,表现出极高的运行效率。

原文摘要 · Abstract (English)

We introduce O-EENC-SD: an end-to-end online speaker diarization system based on EEND-EDA, featuring a novel RNN-based stitching mechanism for online prediction. In particular, we develop a novel centroid refinement decoder whose usefulness is assessed through a rigorous ablation study. Our system provides key advantages over existing methods: a hyperparameter-free solution compared to unsupervised clustering approaches, and a more efficient alternative to current online end-to-end methods, which are computationally costly. We demonstrate that O-EENC-SD is competitive with the state of the art in the two-speaker conversational telephone speech domain, as tested on the CallHome dataset. Our results show that O-EENC-SD provides a great trade-off between DER and complexity, even when working on independent chunks with no overlap, making the system extremely efficient.

说话人分离端到端在线处理高效模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。