arXiv:2608.01423cs.AIcs.GT2026-08

提出评估文本生成的双重视角,兼顾人类评分相关性与抗作弊能力。

Scoring Rules! Statistical and Strategic Alignment for Text Evaluation Metrics

  • 设计三原则:相关性、降质敏感性、抗操纵性,评估指标可靠性。
  • 基于互信息的新指标在保持高相关性的同时显著提升抗操纵性。
  • 适合关注评估公平性与模型优化可信度的研究者使用。

基于参考文本的文本评估指标通过对比候选回答与参考答案来打分,其可靠性通常以与人类评分的相关性衡量。但随着这些指标被用作优化目标,仅靠相关性已不足:智能体可能策略性地操纵得分。本文从统计对齐和战略对齐两个互补角度研究该问题。统计对齐指指标与人类评分高度相关;战略对齐指指标能抵抗不增加任务相关信息的扰动。本文提出三项测试原则:人类评分相关性、退化敏感性、操纵鲁棒性,分别评估指标是否符合人类判断、是否惩罚低效信息丢失、是否抵御策略性得分膨胀。进一步构建基于互信息的统一设计框架,将现有及新指标分解为四个选择:信息度量方式、估计方法、文本表示、预测机制。在同行评审、摘要生成和问答任务中发现,大语言模型作为评判者虽具高相关性,却易被操纵;而基于互信息的指标显著提升抗操纵性。该框架还揭示了一种新指标,在实验中实现最强整体鲁棒性,同时保持竞争力的人类评分相关性。

原文摘要 · Abstract (English)

Reference-based text evaluation metrics, which are widely used to assess natural language generation systems, score a candidate response by comparing it with a reference response. The reliability of an evaluation metric is usually judged by its statistical correlation with human ratings. However, as these metrics are increasingly used as optimization objectives, correlation alone is no longer sufficient: agents may strategically game the evaluation metric. We study this issue through two complementary notions of alignment. A metric is statistically aligned if it correlates with human ratings and strategically aligned if it resists perturbations that do not add task-relevant information. We make two contributions. First, we propose test principles for reference-based metrics consisting of human-rating correlation, degradation sensitivity, and manipulation robustness. These principles evaluate whether a metric agrees with human judgments, penalizes low-effort information loss, and resists strategic score inflation. Second, we develop a unified design framework for mutual-information-based metrics that decomposes existing and new metrics into four choices: information measure, estimation method, text representation, and prediction mechanism. Across peer review, summarization, and question answering, we find that strong human-rating correlation does not imply strategic alignment: LLM-as-a-Judge achieves high correlation but is susceptible to manipulation. In contrast, mutual-information-based metrics substantially improve manipulation robustness. Our framework also uncovers a new metric that achieves the strongest overall robustness in our experiments while remaining competitive on human-rating correlation.

文本评估互信息抗操纵指标设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。