arXiv:2410.20564cs.HCcs.SD2024-10

用语音速度变化提升听觉无障碍的语音识别错误检测能力

Using Confidence Scores to Improve Eyes-free Detection of Speech Recognition Errors

  • 根据识别置信度动态调整语音播报速度,低置信时减速
  • 误差检测率提升12%,决策时间缩短11%
  • 适合盲人或视障用户在无视觉反馈场景使用

对话系统严重依赖语音识别来理解并响应用户指令。尽管语音识别准确率已有提升,但错误仍可能发生,显著影响用户体验。虽然视觉反馈有助于发现错误,但在盲人或低视力用户中可能不实用。本研究探讨通过根据识别器对结果的置信度调节转录文本的音频输出,来改善错误检测。结果显示,在识别器表现出不确定性时选择性地放慢语音,使参与者检测错误的能力相对提升了12%,同时将听取识别结果并判断是否有错误所需时间减少了11%。

原文摘要 · Abstract (English)

Conversational systems rely heavily on speech recognition to interpret and respond to user commands and queries. Despite progress on speech recognition accuracy, errors may still sometimes occur and can significantly affect the end-user utility of such systems. While visual feedback can help detect errors, it may not always be practical, especially for people who are blind or low-vision. In this study, we investigate ways to improve error detection by manipulating the audio output of the transcribed text based on the recognizer's confidence level in its result. Our findings show that selectively slowing down the audio when the recognizer exhibited uncertainty led to a 12% relative increase in participants' ability to detect errors compared to uniformly slowing the audio. It also reduced the time it took participants to listen to the recognition result and decide if there was an error by 11%.

语音识别无障碍设计置信度人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。