提出行为正确性假设,揭示评估模型在可控条件下的真实表现差异。
Beyond Aggregate Scores: Behavioral Correctness Assumptions for Assessing Reference-Based Automatic Evaluation Methods

- 通过可控文本变换定义评估模型的行为预期
- 发现不同评估方法在行为上存在显著差异,无一全满足假设
- 适合关注评估系统内在稳定性和可解释性的研究者
自动参考基准评估方法在自然语言生成系统评估中起关键作用。现有元评估主要衡量与人类判断或基准标签的一致性,难以揭示评估器在受控条件下的行为特征。本文提出行为正确性假设框架,定义了保持正确性与改变正确性的假设分类,并通过可控响应变换实现其操作化,明确预期评分行为。我们评估了多种词法、字符级、语义、基于大模型及混合型评估器,分析其在假设层面的行为、稳定性、敏感性、重复运行变异性、配置敏感性与可复现性。实验显示不同评估范式存在明显行为权衡:无一种评估器满足所有假设;聚合性能相近的评估器可能表现出截然不同的行为特征。结果表明,行为正确性假设能揭示传统聚合元评估所掩盖的诊断信息。
原文摘要 · Abstract (English)
Automated reference-based evaluation methods play a critical role in assessing natural language generation systems. Existing meta-evaluation primarily measures agreement with human judgments or benchmark labels, providing limited insight into evaluator behavior under controlled conditions. We introduce behavioral correctness assumptions, a complementary framework for evaluating reference-based automatic evaluation methods. We define a taxonomy of correctness-preserving and correctness-altering assumptions and operationalize them through controlled response transformations that specify expected scoring behaviors. We evaluate diverse lexical, character-level, semantic, LLM-based, and hybrid evaluators and analyze their assumption-level behavior, stability, sensitivity, repeat-run variability, configuration sensitivity, and reproducibility. Our experiments reveal distinct behavioral trade-offs across evaluation paradigms: no evaluator satisfies all proposed correctness assumptions, and evaluators with similar aggregate performance can exhibit substantially different behavioral profiles. These findings demonstrate that behavioral correctness assumptions provide diagnostic information obscured by conventional aggregate meta-evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。