arXiv:2606.16009cs.CLcs.HC2026-06

让机器翻译更像真人,提升实时对话的流畅性与可靠性

Bridging the Usability Gap: Lessons from Interpreting Studies for Machine Interpreting Design

  • 从口译研究中提炼出三类关键设计原则:主动响应、情境感知和交互自适应
  • 现有系统虽文本准确率高,但实际沟通中常因缺乏上下文理解而失败
  • 适合语音翻译、人机交互及多语言通信系统研发者参考

机器口译(MI)作为语音翻译的实时应用,已在标准基准上取得显著进展,部分系统在文本忠实度上接近人类水平。然而用户体验仍远逊于人工口译,呈现出我们称之为‘准确性幻觉’的现象:系统在指标上看似准确,实际却难以支持流畅、目标导向的交互。本文将MI界定为语音翻译的一个独立子领域,强调其独特性及需要基于交际有效性而非孤立忠实度指标的评估方法。基于口译研究,我们识别出当前系统忽视的专业口译核心维度,并归纳出三个相互关联的设计优先级:代理能力(上下文敏感的主动介入与纠错)、情境对齐(多模态与话语级情境感知)、用户体验(通过真实交互实现自适应优化)。这三大原则共同指明了缩小可用性差距、实现真正实时多语言沟通的路径。

原文摘要 · Abstract (English)

Machine interpreting (MI), the live, real-time application of speech translation, has achieved remarkable progress on standard benchmarks, with some systems approaching human parity on textual fidelity. Yet the user experience remains far inferior to interpreter-mediated communication, revealing what we term the accuracy illusion: systems that appear accurate on paper but fail in practice to support smooth, goal-oriented interaction. This paper defines MI as a distinct subfield of speech translation, with its own characteristics and the need for evaluation methods grounded in communicative effectiveness rather than isolated fidelity metrics. Drawing on insights from interpreting studies, we identify critical dimensions of professional interpreting practice that are overlooked by current systems, and consolidate them into three interdependent design priorities for future MI: agency (context-sensitive initiative and repair), grounding (multimodal and discourse-level situational awareness), and experience (adaptive improvement through real interaction). Together, these priorities chart a path toward closing the usability gap and enabling systems that can sustain authentic multilingual communication in real time.

机器口译人机交互语音翻译用户体验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。