arXiv:2508.02112eess.AS2025-08中稿 · IEEE Transactions …被引 11

针对多人长语音识别,提出更准确的错误率评估方法。

Word Error Rate Definitions and Algorithms for Long-Form Multi-talker Speech Recognition

  • 统一梳理现有错误率定义,明确适用场景。
  • 提出新指标DI-cpWER,量化说话人混淆带来的误差。
  • 设计高效算法,使复杂指标计算更快且误差低于0.1%。

主流语音识别评估指标词错误率(WER)在处理长时多说话人语音识别结果时面临挑战。现有方法分为基于说话人归属的cpWER、tcpWER与忽略说话人归属的ORC-WER、MIMO-WER,它们衡量不同类型的错误(如时间错位)。本文首次系统对比这些度量标准,并提出改进方案:引入去说话人归属的cpWER(DI-cpWER),其与cpWER的差异可反映说话人混淆对总错误率的影响。为辅助人工判断,提出可视化参考与假设序列对齐的方法。针对部分指标计算复杂度高问题,设计贪心近似算法,使ORC-WER与DI-cpWER实现多项式复杂度,实验显示误差低于0.1%。同时将tcpWER的时间约束融入ORC-WER与MIMO-WER,显著降低计算开销。

原文摘要 · Abstract (English)

The predominant metric for evaluating speech recognizers, the Word Error Rate (WER) has been extended in different ways to handle transcripts produced by long-form multi-talker speech recognizers. These systems process long transcripts containing multiple speakers and complex speaking patterns so that the classical WER cannot be applied. There are speaker-attributed approaches that count speaker confusion errors, such as the concatenated minimum-permutation WER cpWER and the time-constrained cpWER (tcpWER), and speaker-agnostic approaches, which aim to ignore speaker confusion errors, such as the Optimal Reference Combination WER (ORC-WER) and the MIMO-WER. These WERs evaluate different aspects and error types (e.g., temporal misalignment). A detailed comparison has not been made. We therefore present a unified description of the existing WERs and highlight when to use which metric. To further analyze how many errors are caused by speaker confusion, we propose the Diarization-invariant cpWER (DI-cpWER). It ignores speaker attribution errors and its difference to cpWER reflects the impact of speaker confusions on the WER. Since error types cannot reliably be classified automatically, we discuss ways to visualize sequence alignments between the reference and hypothesis transcripts to facilitate the spotting of errors by a human judge. Since some WER definitions have high computational complexity, we introduce a greedy algorithm to approximate the ORC-WER and DI-cpWER with high precision ($<0.1\%$ deviation in our experiments) and polynomial complexity instead of exponential. To improve the plausibility of the metrics, we also incorporate the time constraint from the tcpWER into ORC-WER and MIMO-WER, also significantly reducing the computational complexity.

语音识别错误率评估多说话人算法优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。