arXiv:2505.16281cs.CL2025-05EMNLP被引 4

用分层多智能体框架提升机器翻译评估的细粒度与准确性

HiMATE: A Hierarchical Multi-Agent Framework for Machine Translation Evaluation

  • 基于MQM错误类型构建分层多智能体系统,逐级分析翻译错误
  • 在错误定位和严重性判断上比最优基线提升89%的F1分数
  • 适合需要高精度翻译质量评估的研究者与工业应用

大型语言模型(LLMs)的发展使自动评估更具灵活性和可解释性。在机器翻译评估领域,基于多维质量度量(MQM)错误标注的LLM方法能生成更贴近人类判断的结果。然而,现有基于LLM的评估方法仍难以准确识别错误片段并评估其严重程度。本文提出HiMATE——一种面向机器翻译评估的分层多智能体框架。我们认为,现有方法未能充分挖掘MQM层级结构中的细粒度语义信息。为此,我们构建了一个基于MQM错误类型的分层多智能体系统,实现子类型错误的精细化评估。通过引入模型自我反思能力及异构信息下智能体间的协作讨论,有效缓解了系统性幻觉问题。实验表明,HiMATE在多个数据集上均优于现有基线,尤其在错误片段检测与严重性评估方面表现显著,平均F1分数较最优基线提升89%。代码与数据已公开于https://github.com/nlp2ct-shijie/HiMATE。

原文摘要 · Abstract (English)

The advancement of Large Language Models (LLMs) enables flexible and interpretable automatic evaluations. In the field of machine translation evaluation, utilizing LLMs with translation error annotations based on Multidimensional Quality Metrics (MQM) yields more human-aligned judgments. However, current LLM-based evaluation methods still face challenges in accurately identifying error spans and assessing their severity. In this paper, we propose HiMATE, a Hierarchical Multi-Agent Framework for Machine Translation Evaluation. We argue that existing approaches inadequately exploit the fine-grained structural and semantic information within the MQM hierarchy. To address this, we develop a hierarchical multi-agent system grounded in the MQM error typology, enabling granular evaluation of subtype errors. Two key strategies are incorporated to further mitigate systemic hallucinations within the framework: the utilization of the model's self-reflection capability and the facilitation of agent discussion involving asymmetric information. Empirically, HiMATE outperforms competitive baselines across different datasets in conducting human-aligned evaluations. Further analyses underscore its significant advantage in error span detection and severity assessment, achieving an average F1-score improvement of 89% over the best-performing baseline. We make our code and data publicly available at https://github.com/nlp2ct-shijie/HiMATE.

机器翻译评估框架多智能体大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。