arXiv:2507.01001cs.CLcs.AI2025-07NeurIPS被引 18

让科研人员投票比拼大模型的科学文献问答能力

SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks

  • 采用社区投票方式评估大模型在科学文献任务中的表现
  • 已收集超2万条研究人员投票,覆盖47个模型
  • 适合关注自动化评测与科学AI评估的研究者

我们提出SciArena,一个开放协作的平台,用于评估基础模型在科学文献依赖型任务上的表现。与传统基准不同,SciArena借鉴聊天机器人竞技场模式,让研究者直接参与模型对比投票。通过集体智慧,平台实现对需基于文献、生成长文本的开放式科学任务的社区驱动评估。目前支持47个基础模型,已收集来自多个科学领域的超2万条人类研究员投票。数据分析显示数据质量高。我们基于排行榜讨论了模型表现结果。为推动文献任务自动化评估系统研究,我们发布SciArena-Eval,一个基于收集偏好数据的元评估基准,通过比较模型对答案质量的两两判断与人工投票的一致性,衡量其评估准确性。实验揭示该基准的挑战,凸显亟需更可靠的自动评估方法。

原文摘要 · Abstract (English)

We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons. By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses. The platform currently supports 47 foundation models and has collected over 20,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality. We discuss the results and insights based on the model ranking leaderboard. To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on collected preference data. It measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark's challenges and emphasize the need for more reliable automated evaluation methods.

科学智能模型评估社区投票文献理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。