arXiv:2501.16641eess.AScs.HC2025-01中稿 · ICASSP 2025被引 1

实时会议中精准识别多人发言,准确率提升超50%。

SCDiar: a streaming diarization system based on speaker change detection and speech recognition

  • 按语音片段分段,通过说话人切换检测定位发言区间
  • 在十人以上会议中准确率比旧系统高53.6%
  • 适合需要实时多人对话分析的会议场景

在长达数小时的会议场景中,实时语音流常因说话人分离不准确导致识别错误和人数统计偏差。为此,我们提出SCDiar,一个基于说话人切换检测(SCD)模块在词级别分割语音片段的流式语音分离系统。基于这些片段,引入多项优化策略,高效为每位说话人选择最佳语音段。该方法在多个基准测试中表现显著提升。尤其在包含超过十名参与者的实际会议数据上,准确率相比以往系统最高提升53.6%,大幅缩小了在线与离线系统的性能差距。

原文摘要 · Abstract (English)

In hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count errors. To address this challenge, we propose SCDiar, a system that operates on speech segments, split at the token level by a speaker change detection (SCD) module. Building on these segments, we introduce several enhancements to efficiently select the best available segment for each speaker. These improvements lead to significant gains across various benchmarks. Notably, on real-world meeting data involving more than ten participants, SCDiar outperforms previous systems by up to 53.6\% in accuracy, substantially narrowing the performance gap between online and offline systems.

语音分离实时处理会议分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。