arXiv:2604.11246cs.CL2026-04

为长篇生成答案设计了更贴近人类评分的评估框架。

Judge Like Human Examiners: A Weighted Importance Multi-Point Evaluation Framework for Generative Tasks with Long-form Answers

论文配图:Judge Like Human Examiners: A Weighted Importance Multi-Point Evaluation Framework for Generative Tasks with Long-form Answers
图 1 · 摘自论文原文
  • 将参考答案拆解为加权的上下文相关评分点,实现细粒度评估。
  • 在10个生成任务上与人工评分相关性更高,尤其擅长识别内容真伪。
  • 适合需要精准质量评估的学术写作、问答系统等场景。

生成任务中长篇回答的质量评估仍具挑战性,因理想答案通常包含多个语义独立但互补的要素,需进行细粒度分解评估。现有方法依赖任务级评分标准或问题感知检查清单,但仍存在两个不足:1)难以判断回答是否真实基于给定上下文;2)无法捕捉参考答案各部分的重要性差异。受人类评分员启发,我们提出加权重要性多点评估(WIMPE)框架,将每个参考答案分解为加权的上下文绑定评分点。设计了两种互补指标:加权逐点对齐(WPA)衡量模型回答与参考答案的契合度,逐点冲突惩罚(PCP)检测矛盾之处。在10个生成任务上的大量实验表明,WIMPE与人工标注的相关性显著更高。

原文摘要 · Abstract (English)

Evaluating the quality of model responses remains challenging in generative tasks with long-form answers, as the expected answers usually contain multiple semantically distinct yet complementary factors that should be factorized for fine-grained assessment. Recent evaluation methods resort to relying on either task-level rubrics or question-aware checklists. However, they still 1) struggle to assess whether a response is genuinely grounded in provided contexts; 2) fail to capture the heterogeneous importance of different aspects of reference answers. Inspired by human examiners, we propose a Weighted Importance Multi-Point Evaluation (WIMPE) framework, which factorizes each reference answer into weighted context-bound scoring points. Two complementary metrics, namely Weighted Point-wise Alignment (WPA) and Point-wise Conflict Penalty (PCP), are designed to measure the alignment and contradiction between model responses and reference answers. Extensive experiments on 10 generative tasks demonstrate that WIMPE achieves higher correlations with human annotations.

生成评估多点评分人类对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。