arXiv:2507.16136cs.SDcs.AI2025-07被引 4

构建首个统一评测框架,系统化对比语音分割模型性能

SDBench: A Comprehensive Benchmark Suite for Speaker Diarization

  • 整合13个数据集,提供标准化评估流程
  • 实测6大主流系统,发现精度与速度显著权衡
  • 助力新模型快速验证,提升开发效率

当前先进的语音分割系统在不同数据集上误差率差异显著,且系统间比较需严格遵循统一的数据划分和指标定义才能实现公平对比。我们提出SDBench(语音分割基准),一个开源基准套件,集成13个多样化数据集,并内置工具支持对端侧与服务器端系统的细粒度性能分析。SDBench支持可复现的评估与长期新系统集成。为验证其有效性,我们基于Pyannote v3构建了聚焦推理效率的SpeakerKit,借助SDBench快速完成消融实验,使其推理速度比Pyannote v3快9.6倍,同时保持相近的错误率。我们对包括Deepgram、AWS Transcribe和Pyannote AI API在内的6个前沿系统进行了基准测试,揭示了准确率与速度之间的关键权衡。

原文摘要 · Abstract (English)

Even state-of-the-art speaker diarization systems exhibit high variance in error rates across different datasets, representing numerous use cases and domains. Furthermore, comparing across systems requires careful application of best practices such as dataset splits and metric definitions to allow for apples-to-apples comparison. We propose SDBench (Speaker Diarization Benchmark), an open-source benchmark suite that integrates 13 diverse datasets with built-in tooling for consistent and fine-grained analysis of speaker diarization performance for various on-device and server-side systems. SDBench enables reproducible evaluation and easy integration of new systems over time. To demonstrate the efficacy of SDBench, we built SpeakerKit, an inference efficiency-focused system built on top of Pyannote v3. SDBench enabled rapid execution of ablation studies that led to SpeakerKit being 9.6x faster than Pyannote v3 while achieving comparable error rates. We benchmark 6 state-of-the-art systems including Deepgram, AWS Transcribe, and Pyannote AI API, revealing important trade-offs between accuracy and speed.

语音分割基准测试效率优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。