按严重程度分类法评估法语语音识别错误,更贴近人耳理解实际。
A Benchmark of French ASR Systems Based on Error Severity
- 基于内容词和语境划分错误严重等级,共四类细分。
- 在10个顶尖法语ASR系统上测试,揭示各系统优劣差异。
- 适合关注用户体验的语音识别研究者与产品设计者。
自动语音识别(ASR)转录错误通常用与参考文本对比的指标评估,如词错误率(WER),仅衡量拼写偏差或基于语义得分的指标。然而,这些方法常忽略人类对错误可理解性的感知。为此,本文提出一种新评估方法,依据客观语言学标准、上下文模式及以内容词为分析单位,将错误分为四个严重等级并细分为子类型。该指标被应用于涵盖传统HMM模型与端到端模型的10个先进法语ASR系统的基准测试。结果揭示各系统在错误分布与用户阅读体验上的差异,识别出最利于用户理解的系统。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) transcription errors are commonly assessed using metrics that compare them with a reference transcription, such as Word Error Rate (WER), which measures spelling deviations from the reference, or semantic score-based metrics. However, these approaches often overlook what is understandable to humans when interpreting transcription errors. To address this limitation, a new evaluation is proposed that categorizes errors into four levels of severity, further divided into subtypes, based on objective linguistic criteria, contextual patterns, and the use of content words as the unit of analysis. This metric is applied to a benchmark of 10 state-of-the-art ASR systems on French language, encompassing both HMM-based and end-to-end models. Our findings reveal the strengths and weaknesses of each system, identifying those that provide the most comfortable reading experience for users.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。