研究长文本问答评估方法,发现人类偏好判断只适合系统级对比。
Deep Research, Shallow Evaluation: A Case Study in Meta-Evaluation for Long-Form QA Benchmarks
- 用科学领域问答数据集进行多轮人工对比评估
- 发现人类偏好对细粒度指标评估不可靠,易受主观影响
- 建议未来评估应匹配专家评委与具体评估目标
近期长文本生成系统广泛可用,推动了基于大模型评分和事实验证的评估框架发展,以及旨在验证这些方法的元评估框架。现有元评估常通过比较评估结果与人类成对偏好来估计评估质量。然而,已有研究表明人类成对偏好可能过于简单,难以捕捉专家期望的细微差别。本文以面向科学领域检索增强型深度研究问答的ScholarQA-CS2基准为案例,通过人类成对偏好判断全面验证该基准,深入分析该方法的优势、局限与混淆因素。研究显示,成对偏好排名适用于系统层面评估,而可靠的关键指标评估需依赖显式打分和专家标注,主观性仍是核心挑战。基于此,本文提出未来元评估设计的实用指南,强调评估方法、标注者专业性和报告实践的匹配。通过揭示这些方法论问题,旨在提升深度研究系统评估标准。
原文摘要 · Abstract (English)
Recent advances have made long-form report-generating systems widely available. This has prompted evaluation frameworks that use LLM-as-judge protocols and claim verification, along with meta-evaluation frameworks that seek to validate these methods. Many of the meta-evaluations estimate an evaluation quality's by comparing its assessments against human pairwise preferences. Prior work, however, suggests that human pairwise preference may be overly simplistic and can fail to capture nuances of expert expectations. We conduct a case study in meta-evaluation for long-form QA benchmarks using ScholarQA-CS2, a benchmark designed for assessing retrieval-augmented deep-research QA in the scientific domain. We comprehensively validate the benchmark through human pairwise preference judgments, then critically examine the strengths, weaknesses, and confounders of this approach. We show that pairwise preference rankings are best suited for system-level evaluation, while explicit metric-wise annotations and expert annotators are critical for reliable metric-level assessment, with subjectivity remaining a key challenge. Based on our findings, we offer practical guidelines for designing future meta-evaluations that better align evaluation methods, annotator expertise, and reporting practices. By surfacing these methodological challenges, we aim to advance evaluation standards for deep-research systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。