arXiv:2602.23603cs.CLcs.AI2026-02中稿 · ed

构建130万条人工偏好数据,提升长文本问答评估可信度

LFQA-HP-1M: A Large-Scale Human Preference Dataset for Long-Form Question Answering

  • 基于9个评价维度构建人工偏好标注集
  • 线性模型性能媲美顶尖大模型评估器
  • 揭示大模型评估的偏见与脆弱性,适合评测研究者

长文本问答(LFQA)需要对多句解释性回答进行细致评估,但现有指标常无法反映人类判断。本文提出LFQA-HP-1M,一个包含130万条人工成对偏好标注的大规模数据集。我们设计了9项答案质量评价维度,并证明基于这些特征的简单线性模型表现可媲美当前最先进的大语言模型评估器。进一步分析显示,大模型评估器存在传递性不一致、位置偏差和冗长性偏见,且易受对抗性扰动影响。本工作提供了目前最大的公开LFQA偏好数据集,以及一套透明可靠的评估框架。

原文摘要 · Abstract (English)

Long-form question answering (LFQA) demands nuanced evaluation of multi-sentence explanatory responses, yet existing metrics often fail to reflect human judgment. We present LFQA-HP-1M, a large-scale dataset comprising 1.3M human pairwise preference annotations for LFQA. We propose nine rubrics for answer quality evaluation, and show that simple linear models based on these features perform comparably to state-of-the-art LLM evaluators. We further examine transitivity consistency, positional bias, and verbosity biases in LLM evaluators and demonstrate their vulnerability to adversarial perturbations. Overall, this work provides one of the largest public LFQA preference datasets and a rubric-driven framework for transparent and reliable evaluation.

问答评估偏好数据大模型评测人工标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。