构建首个法语人工标注语音识别评估数据集,验证不同指标与人类感知的一致性。
HATS: An Open data set Integrating Human Perception Applied to the Evaluation of Automatic Speech Recognition Metrics

- 收集143人对自动语音识别结果的偏好选择,建立法语人工感知数据集。
- 发现嵌入式指标(如BERTscore)比传统词错误率更接近人类判断。
- 为语音识别评估提供基于真实人类感知的基准,适合评测系统优化研究者。
传统自动语音识别(ASR)系统评估依赖词错误率(WER),但该指标难以全面反映系统表现。为此,本文提出人类感知转录对比数据集HATS,包含143名参与者对多个ASR系统生成的双候选转录进行主观偏好选择。研究分析了词汇型与嵌入式评估指标(如BERTscore、语义距离)与人类判断的相关性,结果显示嵌入式指标在捕捉人类感知方面优于传统方法。该数据集为改进语音识别评估提供了可量化的基准。
原文摘要 · Abstract (English)
Conventionally, Automatic Speech Recognition (ASR) systems are evaluated on their ability to correctly recognize each word contained in a speech signal. In this context, the word error rate (WER) metric is the reference for evaluating speech transcripts. Several studies have shown that this measure is too limited to correctly evaluate an ASR system, which has led to the proposal of other variants of metrics (weighted WER, BERTscore, semantic distance, etc.). However, they remain system-oriented, even when transcripts are intended for humans. In this paper, we firstly present Human Assessed Transcription Side-by-side (HATS), an original French manually annotated data set in terms of human perception of transcription errors produced by various ASR systems. 143 humans were asked to choose the best automatic transcription out of two hypotheses. We investigated the relationship between human preferences and various ASR evaluation metrics, including lexical and embedding-based ones, the latter being those that correlate supposedly the most with human perception.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。