arXiv:2504.11582cs.CL2025-04ACL被引 9

让不懂法语的人也能判断机器翻译质量,靠问答自动检测关键错误。

AskQE: Question Answering as Automatic Evaluation for Machine Translation

  • 用问答框架生成问题,基于前提事实引导提问来发现翻译错误。
  • 在生物医学数据集上,与人工评分相关性更高,决策准确率达87.3%。
  • 适合无目标语言能力的编辑或审校人员快速评估翻译可信度。

如何让不懂法语的英语使用者判断法语机器翻译是否足够好以共享?现有机器翻译错误检测和质量评估方法无法应对这一实际场景。我们提出AskQE,一种基于问题生成与回答的框架,用于检测关键翻译错误并提供可操作反馈,使用户即使不具备目标语言知识也能决定是否接受或拒绝机器翻译结果。利用ContraTICO数据集(新冠领域对比性合成翻译错误),我们探索了AskQE的设计选择,并开发出基于LLaMA-3 70B和前提事实引导的问题生成优化版本。在自然发生的机器翻译错误数据集BioMQM上的评估显示,AskQE相较于其他质量评估指标,在人类评分上的肯德尔等级相关系数更高,决策准确率达到87.3%。

原文摘要 · Abstract (English)

How can a monolingual English speaker determine whether an automatic translation in French is good enough to be shared? Existing MT error detection and quality estimation (QE) techniques do not address this practical scenario. We introduce AskQE, a question generation and answering framework designed to detect critical MT errors and provide actionable feedback, helping users decide whether to accept or reject MT outputs even without the knowledge of the target language. Using ContraTICO, a dataset of contrastive synthetic MT errors in the COVID-19 domain, we explore design choices for AskQE and develop an optimized version relying on LLaMA-3 70B and entailed facts to guide question generation. We evaluate the resulting system on the BioMQM dataset of naturally occurring MT errors, where AskQE has higher Kendall's Tau correlation and decision accuracy with human ratings compared to other QE metrics.

机器翻译质量评估问答系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。