为不同翻译样本动态匹配评估标准,提升细粒度质量评估准确率。
Rubric-as-Experts: Case-Specific MQM Rubrics for Translation Quality Evaluation

- 基于预定义的MQM分类体系,按需动态选择评估细粒度和子类型空间。
- 在多个模型规模下,相比固定规则,提升了MCC指标并减少误报。
- 适合需要精准定位翻译错误的场景,如机器翻译质量评测系统。
大语言模型在细粒度翻译质量评估(QE)中展现出强大潜力,但现有基于MQM的方法通常采用统一的评估规则,无法适应不同翻译实例在错误复杂度、歧义性和评估粒度上的差异。我们发现更大的MQM子类型空间虽提升错误覆盖范围,却也增加误报;而不同翻译样本对评估粒度的需求各异,表明应为每例动态分配评估空间。为此,我们提出一种案例特定的动态评估框架,在保持预设MQM分类体系的前提下,自适应地为每个翻译实例选择合适的子类型集合与评估粒度。在多个模型规模下的WMT细粒度评估基准测试中,该方法显著提升MCC值,并实现更清晰的错误定位。结果表明,结合结构化MQM规则与案例自适应分配,是提升基于大模型翻译评估效果的有效策略。
原文摘要 · Abstract (English)
Large language models (LLMs) have shown strong potential in fine-grained translation quality evaluation (QE), yet existing MQM-based approaches typically rely on fixed rubric configurations shared across all translation samples. However, translation instances often differ substantially in error complexity, ambiguity, and required evaluation granularity, making static rubric allocation suboptimal for span-level error detection. We find that larger MQM subtype spaces improve error coverage but also introduce more false positives, while different translation instances prefer different rubric granularities, suggesting that evaluation spaces should be allocated dynamically for each case. Motivated by these observations, we propose a case-specific dynamic rubric framework that adaptively constructs MQM evaluation spaces for individual translation instances. Unlike fully free-form rubric generation methods, our framework remains grounded in the predefined MQM taxonomy while dynamically selecting suitable subtype spaces and evaluation granularity for different cases. Experiments on WMT span-level QE benchmarks across multiple model scales demonstrate that the proposed framework consistently improves MCC and produces cleaner span-level error localization compared with static rubric settings. Our results suggest that combining structured MQM rubrics with case-specific adaptive allocation is an effective strategy for fine-grained LLM-based translation evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。