用75个领域调查文章构建大规模科研问答评测集
ResearchQA: Evaluating Scholarly Question Answering at Scale Across 75 Fields with Survey-Mined Questions and Rubrics
- 从75个领域的综述文章提取2.1万道问题和16万条评分标准
- 现有大模型平均仅覆盖75%评分项,引用类要求满足不足11%
- 适合做跨学科科研问答评估,尤其关注答案完整性与严谨性
评估长篇科研问答回答严重依赖专家标注,限制在人工智能等可动员同行的领域。但科研知识广泛分布于综述文章中。我们提出ResearchQA,通过提炼75个研究领域的综述文章,生成21,000个问题和160,000条评分标准。问题与评分标准共同源自综述章节,评分项包括引用文献、解释说明、描述局限性等。8个领域的31名博士级标注者评估显示,90%的问题符合博士级信息需求,87%的评分项需至少一个句子以上强调。我们利用ResearchQA对18个系统进行7,600次对比评测。所测参数化或检索增强系统均未超过70%的评分项覆盖,表现最佳系统达75%。错误分析显示,该系统仅完全满足不足11%的引用评分项、48%的局限性项和49%的比较项。数据已公开,以支持更全面的多领域评估。
原文摘要 · Abstract (English)
Evaluating long-form responses to research queries heavily relies on expert annotators, restricting attention to areas like AI where researchers can conveniently enlist colleagues. Yet, research expertise is abundant: survey articles consolidate knowledge spread across the literature. We introduce ResearchQA, a resource for evaluating LLM systems by distilling survey articles from 75 research fields into 21K queries and 160K rubric items. Queries and rubrics are jointly derived from survey sections, where rubric items list query-specific answer evaluation criteria, i.e., citing papers, making explanations, and describing limitations. 31 Ph.D. annotators in 8 fields judge that 90% of queries reflect Ph.D. information needs and 87% of rubric items warrant emphasis of a sentence or longer. We leverage ResearchQA to evaluate 18 systems in 7.6K head-to-heads. No parametric or retrieval-augmented system we evaluate exceeds 70% on covering rubric items, and the highest-ranking system shows 75% coverage. Error analysis reveals that the highest-ranking system fully addresses less than 11% of citation rubric items, 48% of limitation items, and 49% of comparison items. We release our data to facilitate more comprehensive multi-field evaluations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。