arXiv:2606.16596cs.CL2026-06

用目标导向场景评估翻译对对话连贯性的影响,发现高质量翻译仍会出错。

How Far Can Machine Translation Quality Take You? Extrinsic Discourse Evaluation in Goal-Oriented Setups

论文配图:How Far Can Machine Translation Quality Take You? Extrinsic Discourse Evaluation in Goal-Oriented Setups
图 1 · 摘自论文原文
  • 设计实体计数任务检测静态对话中的指代一致性
  • 高分翻译系统仍存在指代错误,影响下游任务表现
  • 交互式多智能体游戏揭示长程沟通中的翻译失败

现有机器翻译评估主要基于内在指标,未考察翻译错误对下游任务的影响。本文在静态和交互两种场景下开展外在话语评估:静态场景中,提出实体计数任务作为话语指代一致性的探针;结果表明,高内在翻译质量无法可靠预测下游话语成功,强模型仍会产生指代不一致。交互场景中,采用目标导向的多智能体博弈游戏Welfare Diplomacy作为长程沟通与协作的探针,发现交互特异性翻译错误会影响后续协调效果。研究证明目标导向环境是开展话语敏感型外在翻译评估的有效框架。

原文摘要 · Abstract (English)

Existing machine translation (MT) metrics and discourse-focused evaluations primarily assess translation quality intrinsically, without measuring the downstream consequences of translation errors. In this work, we focus on extrinsic discourse evaluation of machine translation under two distinct regimes: static and interactive. Under the static regime, we propose an entity counting task as a probe of referential consistency in discourse. We show that high intrinsic MT quality does not reliably predict downstream discourse success and strong MT systems still produce referential inconsistencies. For the interactive regime, we study the goal-oriented multi-agent Welfare Diplomacy game as a probe of long-horizon communication and coordination. We find that interaction-specific translation failures impact downstream coordination. Our results highlight goal-oriented environments as a viable framework for discourse-sensitive extrinsic MT evaluation.

机器翻译话语评估多智能体外在评价

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。