通过人类选择性倾听实验,评估语音识别系统在对话中的关键信息捕捉能力。
What Do Humans Hear When Interacting? Experiments on Selective Listening for Evaluating ASR of Spoken Dialogue Systems
- 基于人类听觉聚焦机制设计对话响应实验
- 发现人类能准确识别对话中关键信息片段
- 为语音识别评估提供新的人类行为基准
语音对话系统(SDS)在前端使用自动语音识别(ASR)来捕捉用户话语中的关键信息以生成回应。本文通过对比人类生成对话回应时的转录与参考转录,实验证实了人类在对话中存在选择性倾听现象——即聚焦于重要语义内容的能力。基于此,研究探讨了一种利用人类选择性倾听行为的新式ASR评估方法,可有效识别不同ASR系统与人类在关键信息转录能力上的差距,为更贴近真实交互场景的评估提供可能。
原文摘要 · Abstract (English)
Spoken dialogue systems (SDSs) utilize automatic speech recognition (ASR) at the front end of their pipeline. The role of ASR in SDSs is to recognize information in user speech related to response generation appropriately. Examining selective listening of humans, which refers to the ability to focus on and listen to important parts of a conversation during the speech, will enable us to identify the ASR capabilities required for SDSs and evaluate them. In this study, we experimentally confirmed selective listening when humans generate dialogue responses by comparing human transcriptions for generating dialogue responses and reference transcriptions. Based on our experimental results, we discuss the possibility of a new ASR evaluation method that leverages human selective listening, which can identify the gap between transcription ability between ASR systems and humans.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。