构建细粒度论文理解评测集,揭示大模型在学术阅读上的显著短板
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
- 基于论文审稿回复构建1.5万条人工验证问答对,覆盖科研流程全阶段
- 最强模型GPT-5在精确完整度上仅达68.2%,考虑简洁性后降至37.46%
- 设计可扩展标注框架,支持大模型与人类判断高一致性评估
由于专业科学语篇和复杂图表的存在,基础模型理解研究论文仍具挑战性,而现有评测基准在细粒度评估方面能力有限。为此,我们提出RPC-Bench,一个基于高质量计算机科学论文审稿-回复交互构建的大规模问答评测集,包含15,000条人工验证的问答对。我们设计了与科研流程一致的细粒度分类体系,评估模型在学术语境中回答为何、何为、如何类问题的能力。同时定义了详尽的LLM-人类交互标注框架,支持大规模标注与质量控制。采用LLM-as-a-Judge范式,构建可扩展的评估框架,衡量模型答案的正确性-完整性与简洁性,与人类判断高度一致。实验显示,即使最强模型(GPT-5)在正确完整性上也仅达68.2%,经简洁性调整后降至37.46%,凸显当前模型在精准学术理解上的巨大差距。代码与数据已公开于https://rpc-bench.github.io/。
原文摘要 · Abstract (English)
Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limited fine-grained evaluation at scale. To address this gap, we introduce RPC-Bench, a large-scale question-answering benchmark built from review-rebuttal exchanges of high-quality computer science papers, containing 15K human-verified QA pairs. We design a fine-grained taxonomy aligned with the scientific research flow to assess models' ability to understand and answer why, what, and how questions in scholarly contexts. We also define an elaborate LLM-human interaction annotation framework to support large-scale labeling and quality control. Following the LLM-as-a-Judge paradigm, we develop a scalable framework that evaluates models on correctness-completeness and conciseness, with high agreement to human judgment. Experiments reveal that even the strongest models (GPT-5) achieve only 68.2% correctness-completeness, dropping to 37.46% after conciseness adjustment, highlighting substantial gaps in precise academic paper understanding. Our code and data are available at https://rpc-bench.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。