arXiv:2504.14804cs.CLcs.AI2025-04综述被引 1

梳理文档级翻译自动评估现状与挑战,指明未来改进方向。

Automatic Evaluation Metrics for Document-level Translation: Overview, Challenges and Trends

  • 分类总结无参考、有参考等各类评估方法
  • 指出现有方法依赖句级对齐、缺乏多样性等问题
  • 适合研究机器翻译评估的学者和工程师参考

随着深度学习技术的快速发展,机器翻译领域取得显著进展,特别是大语言模型(LLMs)的出现极大推动了文档级翻译的发展。然而,准确评估文档级翻译质量仍是亟待解决的问题。本文首先介绍文档级翻译的发展现状及其评估的重要性,强调自动评估指标在反映翻译质量与指导系统优化中的关键作用。接着详细分析当前自动评估方案与指标,涵盖有/无参考文本的评估方法,以及传统指标、基于模型的指标和基于LLM的指标。随后探讨当前方法面临的挑战,如参考文本多样性不足、对句级对齐信息的依赖,以及LLM作为裁判方法存在的偏差、不准确和可解释性差等问题。最后展望未来趋势,包括开发更用户友好的文档级评估方法、提升LLM作为裁判的鲁棒性,并提出可能的研究方向,如降低对句级信息的依赖、引入多层级多粒度评估方式,以及训练专门用于机器翻译评估的模型。本研究旨在为文档级翻译的自动评估提供全面分析,揭示未来发展方向。

原文摘要 · Abstract (English)

With the rapid development of deep learning technologies, the field of machine translation has witnessed significant progress, especially with the advent of large language models (LLMs) that have greatly propelled the advancement of document-level translation. However, accurately evaluating the quality of document-level translation remains an urgent issue. This paper first introduces the development status of document-level translation and the importance of evaluation, highlighting the crucial role of automatic evaluation metrics in reflecting translation quality and guiding the improvement of translation systems. It then provides a detailed analysis of the current state of automatic evaluation schemes and metrics, including evaluation methods with and without reference texts, as well as traditional metrics, Model-based metrics and LLM-based metrics. Subsequently, the paper explores the challenges faced by current evaluation methods, such as the lack of reference diversity, dependence on sentence-level alignment information, and the bias, inaccuracy, and lack of interpretability of the LLM-as-a-judge method. Finally, the paper looks ahead to the future trends in evaluation methods, including the development of more user-friendly document-level evaluation methods and more robust LLM-as-a-judge methods, and proposes possible research directions, such as reducing the dependency on sentence-level information, introducing multi-level and multi-granular evaluation approaches, and training models specifically for machine translation evaluation. This study aims to provide a comprehensive analysis of automatic evaluation for document-level translation and offer insights into future developments.

机器翻译评估方法LLM评估文档级翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。