对比12种相关性度量,发现选对方法能显著提升NLG评估结果可靠性。
Analyzing and Evaluating Correlation Measures in NLG Meta-Evaluation
- 分析12种相关性度量在6个数据集上的表现差异
- 全局分组+皮尔逊相关系数效果最好,抗评分粒度干扰强
- 提出判别力、排序一致性和敏感性三维度评估标准
自然语言生成自动评估指标与人工评价之间的相关性常被视为衡量评估指标能力的关键标准。然而,不同的分组方式和相关系数会导致多种相关性度量被用于元评估。在具体评估场景中,以往研究往往直接沿用传统设置,但这些度量之间的特性与差异未受到足够关注。本文基于六个广泛使用的NLG评估数据集和32种评估指标,分析了12种常见相关性度量,发现不同度量确实影响元评估结果。进一步地,我们提出三个反映元评估能力的视角:判别力、排序一致性以及对评分粒度的敏感性。结果表明,采用全局分组与皮尔逊相关系数的度量在判别力和排序一致性上表现最佳;而使用系统级分组或肯德尔相关系数的度量对评分粒度最不敏感。
原文摘要 · Abstract (English)
The correlation between NLG automatic evaluation metrics and human evaluation is often regarded as a critical criterion for assessing the capability of an evaluation metric. However, different grouping methods and correlation coefficients result in various types of correlation measures used in meta-evaluation. In specific evaluation scenarios, prior work often directly follows conventional measure settings, but the characteristics and differences between these measures have not gotten sufficient attention. Therefore, this paper analyzes 12 common correlation measures using a large amount of real-world data from six widely-used NLG evaluation datasets and 32 evaluation metrics, revealing that different measures indeed impact the meta-evaluation results. Furthermore, we propose three perspectives that reflect the capability of meta-evaluation: discriminative power, ranking consistency, and sensitivity to score granularity. We find that the measure using global grouping and Pearson correlation coefficient exhibits the best performance in both discriminative power and ranking consistency. Besides, the measures using system-level grouping or Kendall correlation are the least sensitive to score granularity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。