arXiv:2601.20992cs.CLcs.SD2026-01

改进语音识别评估,支持多参考转录与流式识别分析

asr_eval: Algorithms and tools for multi-reference and streaming speech recognition evaluation

  • 提出支持多参考标注和任意长度插入的字符串对齐算法
  • 在俄语长段落语音上构建新数据集,发现模型会适应标注习惯
  • 提供通用接口,兼容多种离线与流式语音识别模型

我们提出多项语音识别评估改进。首先,设计一种支持多参考标注、任意长度插入及更优词对齐的字符串对齐算法,特别适用于非拉丁字母语言及构词丰富的语言,用于标注嘈杂或长段语音。其次,构建了一个新型长段落真实场景俄语测试集 DiverseSpeech-Ru,采用精细的多参考标注,并对流行的俄语测试集进行多参考重标注,研究其训练集上的微调动态。结果表明,模型常适应数据集特定标注方式,造成评估指标虚高。基于改进的词对齐,开发了流式语音识别评估工具,并可将多条转录对齐以可视化对比。此外,提供多个离线与流式语音识别模型的统一接口。代码将公开。

原文摘要 · Abstract (English)

We propose several improvements to the speech recognition evaluation. First, we propose a string alignment algorithm that supports both multi-reference labeling, arbitrary-length insertions and better word alignment. This is especially useful for non-Latin languages, those with rich word formation, to label cluttered or longform speech. Secondly, we collect a novel test set DiverseSpeech-Ru of longform in-the-wild Russian speech with careful multi-reference labeling. We also perform multi-reference relabeling of popular Russian tests set and study fine-tuning dynamics on its corresponding train set. We demonstrate that the model often adopts to dataset-specific labeling, causing an illusion of metric improvement. Based on the improved word alignment, we develop tools to evaluate streaming speech recognition and to align multiple transcriptions to compare them visually. Additionally, we provide uniform wrappers for many offline and streaming speech recognition models. Our code will be made publicly available.

语音识别评估方法多参考标注流式处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。