评测非字面翻译质量,提出新框架提升评估准确性。
Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
- 设计可复用的反思型评估框架,动态调用专业子代理。
- 在7530条人工标注数据上,相关性提升至少3.2分。
- 适合需精准评估文学、社交平台等复杂文本的场景。
大型语言模型(LLMs)显著推动了机器翻译(MT)发展,尤其在社交媒体、文学等语言复杂领域。此类场景中,翻译常涉及非字面表达,导致传统评估指标失准。为系统检验评估可靠性,我们构建了聚焦非字面翻译的元评估数据集MENT,涵盖四个非字面翻译领域,包含7,530条由人类标注的翻译质量评分,源句与多系统翻译配对。实验发现传统指标和LLM-as-a-Judge均存在缺陷,包括知识截止与评分不一致问题。为此,我们提出RATE框架,核心为具备反思能力的代理,可动态调用专业化子代理。实验表明,RATE在系统级与段落级相关性上相比现有方法至少提升3.2分,且对通用领域评估具有鲁棒性。代码与数据已开源。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have significantly advanced Machine Translation (MT), applying them to linguistically complex domains-such as Social Network Services, literature etc. In these scenarios, translations often require handling non-literal expressions, leading to the inaccuracy of MT metrics. To systematically investigate the reliability of MT metrics, we first curate a meta-evaluation dataset focused on non-literal translations, namely MENT. MENT encompasses four non-literal translation domains and features source sentences paired with translations from diverse MT systems, with 7,530 human-annotated scores on translation quality. Experimental results reveal the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge, particularly the knowledge cutoff and score inconsistency problem. To mitigate these limitations, we propose RATE, a novel agentic translation evaluation framework, centered by a reflective Core Agent that dynamically invokes specialized sub-agents. Experimental results indicate the efficacy of RATE, achieving an improvement of at least 3.2 points in combined system- and segment-level correlation with human judgments compared with current methods. Further experiments demonstrate the robustness of RATE to general-domain MT evaluation. Code and dataset are available at: https://github.com/BITHLP/RATE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。