arXiv:2503.19828cs.CL2025-03NAACL被引 3

评估指标在不同场景下的表现差异,揭示需按上下文选择评价方式。

Contextual Metric Meta-Evaluation by Measuring Local Metric Accuracy

  • 通过局部准确率比较指标,实现针对具体任务的评价
  • 多任务实验显示指标效果随场景变化显著
  • 适合需要精准评估模型性能的研究者和开发者

自动评价指标的元评估——即评估评价指标本身——对准确评测自然语言处理系统至关重要,影响科学探究、模型开发与政策制定。现有方法多关注指标在任意系统输出下的绝对与相对质量,但在实际中,指标常用于高度受限的特定场景,如仅评估某类模型。本文提出一种基于局部指标准确率的上下文化元评估方法。在机器翻译、语音识别与排序任务中,我们发现随着评估上下文变化,各指标的局部准确率在绝对值和相对有效性上均呈现显著差异。这一现象凸显了采用情境化评估取代全局评估的重要性。

原文摘要 · Abstract (English)

Meta-evaluation of automatic evaluation metrics -- assessing evaluation metrics themselves -- is crucial for accurately benchmarking natural language processing systems and has implications for scientific inquiry, production model development, and policy enforcement. While existing approaches to metric meta-evaluation focus on general statements about the absolute and relative quality of metrics across arbitrary system outputs, in practice, metrics are applied in highly contextual settings, often measuring the performance for a highly constrained set of system outputs. For example, we may only be interested in evaluating a specific model or class of models. We introduce a method for contextual metric meta-evaluation by comparing the local metric accuracy of evaluation metrics. Across translation, speech recognition, and ranking tasks, we demonstrate that the local metric accuracies vary both in absolute value and relative effectiveness as we shift across evaluation contexts. This observed variation highlights the importance of adopting context-specific metric evaluations over global ones.

元评估指标评价上下文感知NLP评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。