arXiv:2605.09098cs.CL2026-05ACL

根据源句特点动态调整翻译评估指标权重,提升评估准确性。

Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation

论文配图:Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation
图 1 · 摘自论文原文
  • 基于源句特征动态调节现有评估指标的组合权重。
  • 在多语言对上,基于MLP的组合方法优于线性与高斯过程模型。
  • 软性条件化设计使评估结果更优,适合需要精准评估的场景。

我们提出动态元度量(DMM),一种机器翻译评估框架,通过学习源句特征条件下的现有指标组合来提升评估性能。不同于依赖单一静态集成或语言特异性权重的方法,DMM根据源句片段的属性自适应调整指标组合方式。研究包含硬性条件化(为每个聚类拟合可解释的组合器)和探索性软性条件化扩展(权重随源句聚类责任连续变化)。在多个语言对的WMT评估任务中,使用成对一致性指标在系统级与段级进行评估。结果显示,基于MLP的组合方法在所有设置下均优于线性及高斯过程模型,引入软性条件化进一步提升了性能。

原文摘要 · Abstract (English)

We propose Dynamic Meta-Metrics (DMM), a framework for machine translation evaluation that learns source-sentence conditioned combinations of existing metrics. Rather than relying on a single static ensemble or language-specific weighting, DMM adapts the metric combination based on properties of the source segment. We study hard conditioning, which fits an interpretable combiner per cluster, and an exploratory soft-conditioned extension whose weights vary continuously with source-cluster responsibilities. We evaluate DMM on the WMT Metrics Shared Task data across multiple language pairs using pairwise agreement measures at the system and segment levels. Across settings, MLP-based combinations outperform linear and Gaussian process-based ensembles, and introducing soft conditioning yields gains over linear models.

机器翻译评估方法动态权重

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。