arXiv:2607.11074cs.CL2026-07

评测大模型在科学论文中基于引文回答问题的能力,发现开源模型更高效且准确。

ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers

论文配图:ResearchQA: Benchmarking Citation-Grounded Question-Answering on Scientific Papers
图 1 · 摘自论文原文
  • 构建6211个单篇论文问答对,覆盖8个领域4类问题,强调答案需有引文支持。
  • 开源模型在引文准确性上接近闭源模型,但推理速度快3到6倍。
  • 引入引文匹配器和LLM评分器,使模型评估更清晰可靠,适合科研辅助场景。

大语言模型被越来越多用于辅助科学阅读,但现有评估方法常无法检测答案是否由可验证的引文支持。我们提出ResearchQA,一个包含6,211个单篇论文问答对的基准数据集,涵盖494篇开放获取论文,覆盖八个领域和四种问题类型:查找、理解、多跳与对抗性问题。该基准专为引文支撑评估设计:允许一个结论有多个有效支持段落,并奖励当论文不支持答案时的拒绝回答。我们在引用匹配器和基于LLM的评分器下,对八种主流闭源与开源模型进行评估。结果显示,基于引文的指标比LLM评分更能区分系统表现:章节覆盖率与引文准确率在不同模型间差异显著,而评分器得分则高度集中。此外,开源模型在引文准确率上接近最优闭源模型,同时每例推理延迟降低3至6倍。我们公开了基准、评估工具与评分提示。

原文摘要 · Abstract (English)

Large language models are increasingly used to assist scientific reading, but existing evaluation methods often fail to detect whether answers are supported by verifiable citations. We introduce ResearchQA, a benchmark of 6,211 single-paper question-answer pairs from 494 open-access papers spanning eight domains and four question types: lookup, comprehension, multi-hop, and adversarial. ResearchQA is designed for citation-grounded evaluation: it permits multiple valid supporting passages for a claim and rewards grounded refusal when the source paper does not support an answer. We evaluate eight leading closed- and open-weight models in a citation-grounded chat-with-paper setting using a deterministic citation matcher and an LLM-based rubric evaluator. Citation-based metrics separate systems more clearly than LLM-evaluator scores: section coverage and citation accuracy vary substantially across models, while evaluator scores remain tightly compressed. We further find that open-weight models approach the best closed-model citation accuracy while achieving 3 to 6 times lower per-example latency. We release the benchmark, evaluation harness, and evaluator prompt.

科学问答引文评估大模型评测开源模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。