arXiv:2510.19572eess.AS2025-10被引 6

改进语音分角色识别的聚类阶段,提升复杂场景下的准确性

VBx for End-to-End Neural and Clustering-based Diarization

  • 用过滤不靠谱嵌入+重分配策略优化聚类阶段
  • 在多领域数据集上无需微调即达到顶尖性能
  • 适合需要高鲁棒性、少调参的实用语音系统

本文针对两阶段端到端神经语音分角色识别(EEND-VC)框架中的聚类阶段提出改进。第一阶段采用基于Conformer的EEND模型与WavLM特征,对短窗口内的帧级说话人活动进行推断;第二阶段通过跨窗口聚类说话人嵌入,确定全局说话人身份与数量。本文重点优化第二阶段:剔除短段中不可靠的嵌入并重新分配;同时引入VBx聚类方法,提升在说话人数多、发言时长短等挑战场景下的鲁棒性。在覆盖多个领域的复合基准上评估,未对EEND模型进行微调,也未按数据集调优聚类参数,系统仍表现出良好泛化能力,性能匹配或超过近期最先进水平。

原文摘要 · Abstract (English)

We present improvements to speaker diarization in the two-stage end-to-end neural diarization with vector clustering (EEND-VC) framework. The first stage employs a Conformer-based EEND model with WavLM features to infer frame-level speaker activity within short windows. The identities and counts of global speakers are then derived in the second stage by clustering speaker embeddings across windows. The focus of this work is to improve the second stage; we filter unreliable embeddings from short segments and reassign them after clustering. We also integrate the VBx clustering to improve robustness when the number of speakers is large and individual speaking durations are limited. Evaluation on a compound benchmark spanning multiple domains is conducted without fine-tuning the EEND model or tuning clustering parameters per dataset. Despite this, the system generalizes well and matches or exceeds recent state-of-the-art performance.

语音分角色聚类优化端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。