arXiv:2503.15124cs.HCcs.CL2025-03被引 8

评估语音识别置信度在人工纠错界面中的表现,发现其效果有限。

Evaluating ASR Confidence Scores for Automated Error Detection in User-Assisted Correction Interfaces

  • 基于置信度分数检测错误,但准确率不高
  • 分类器常漏检或误报,错误率仍高
  • 用户认为该方法无帮助,未提升纠错效率

尽管自动语音识别(ASR)技术不断进步,转录错误仍普遍存在,需人工修正。置信度分数可反映ASR结果的可信程度,可能辅助用户发现并纠正错误。本研究通过端到端ASR模型的全面分析及36名参与者的用户实验,评估了置信度分数在错误检测中的可靠性。结果表明,虽然置信度与转录准确率存在相关性,但其错误检测性能受限:分类器频繁遗漏错误或产生大量误报,严重削弱实际应用价值。基于置信度的错误检测既未提高修正效率,也未被参与者视为有用。研究揭示了置信度分数的局限性,强调需发展更先进的方法以增强ASR系统的可解释性与人机交互体验。

原文摘要 · Abstract (English)

Despite advances in Automatic Speech Recognition (ASR), transcription errors persist and require manual correction. Confidence scores, which indicate the certainty of ASR results, could assist users in identifying and correcting errors. This study evaluates the reliability of confidence scores for error detection through a comprehensive analysis of end-to-end ASR models and a user study with 36 participants. The results show that while confidence scores correlate with transcription accuracy, their error detection performance is limited. Classifiers frequently miss errors or generate many false positives, undermining their practical utility. Confidence-based error detection neither improved correction efficiency nor was perceived as helpful by participants. These findings highlight the limitations of confidence scores and the need for more sophisticated approaches to improve user interaction and explainability of ASR results.

语音识别置信度人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。