arXiv:2409.05674cs.SDcs.AI2024-09被引 6

提出新方法评估ASR系统延迟,判断其是否适合实时口译场景。

Assessing Latency in ASR Systems: A Methodological Perspective for Real-Time Use

  • 设计基于用户感知的延迟测量方法,聚焦语音到转录的时间差。
  • 发现现有ASR系统延迟常超实时口译可接受阈值。
  • 适用于外交、医疗等对语言细节敏感的实时场景研究者。

自动语音识别(ASR)系统虽能实现实时转录,但常遗漏人类口译员能捕捉的细微语义。尽管在诸多场景中具有实用价值,口译员——尤其是使用如Dragon等工具的人员——仍提供关键补充,尤其在外交会议等敏感场合,细微语言差异至关重要。口译员不仅察觉这些细节,还能实时调整以提升准确率,而ASR仅处理基础转录任务。然而,ASR系统引入的延迟与实时口译需求不匹配。用户感知的延迟不同于口译延迟,它衡量的是从语音输入到转录输出的时间间隔。为此,本文提出一种新的延迟测量方法,并验证其在实时口译场景中的可用性。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) systems generate real-time transcriptions but often miss nuances that human interpreters capture. While ASR is useful in many contexts, interpreters-who already use ASR tools such as Dragon-add critical value, especially in sensitive settings such as diplomatic meetings where subtle language is key. Human interpreters not only perceive these nuances but can adjust in real time, improving accuracy, while ASR handles basic transcription tasks. However, ASR systems introduce a delay that does not align with real-time interpretation needs. The user-perceived latency of ASR systems differs from that of interpretation because it measures the time between speech and transcription delivery. To address this, we propose a new approach to measuring delay in ASR systems and validate if they are usable in live interpretation scenarios.

ASR延迟评估实时口译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。