arXiv:2512.22899cs.AIcs.CV2025-12被引 2

构建科学智能评估新标准,覆盖从阅读到发现的完整科研流程。

HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery

  • 分五层架构模拟真实科研流程,从理解到创新层层递进。
  • 含8735个跨学科样本,多模态输入支持文本、公式、图表等。
  • 揭示主流模型在发现任务上准确率仅25%,差距显著。

大语言模型和多模态基础模型的快速发展激发了其在科学研究中应用的兴趣。然而,科学智能涵盖从理解基础知识到创造性发现的广泛能力,现有基准仍分散且片面,多聚焦于单一任务,未能反映真实科研的层级性与多学科特性。本文提出 extbf{HiSciBench},一个分层式基准,用于评估基础模型在五个层次上的表现:科学素养(L1)、文献解析(L2)、基于文献的问题回答(L3)、文献综述生成(L4)和科学发现(L5)。该基准包含8,735个精心筛选的实例,覆盖数学、物理、化学、生物、地理和天文六个主要科学领域,支持文本、公式、图像、表格等多模态输入,以及跨语言评估。与以往孤立评估不同,HiSciBench 提供集成、依赖感知的框架,可细致诊断模型在科学推理各阶段的能力。对 GPT-5、DeepSeek-R1 等领先模型的全面评估显示,模型在基础素养任务上最高达69%准确率,但在发现级挑战中骤降至25%。该基准为科学智能评估树立新标准,并为开发更强大、更可靠的模型提供关键洞见。基准将公开发布,以推动后续研究。

原文摘要 · Abstract (English)

The rapid advancement of large language models (LLMs) and multimodal foundation models has sparked growing interest in their potential for scientific research. However, scientific intelligence encompasses a broad spectrum of abilities ranging from understanding fundamental knowledge to conducting creative discovery, and existing benchmarks remain fragmented. Most focus on narrow tasks and fail to reflect the hierarchical and multi-disciplinary nature of real scientific inquiry. We introduce \textbf{HiSciBench}, a hierarchical benchmark designed to evaluate foundation models across five levels that mirror the complete scientific workflow: \textit{Scientific Literacy} (L1), \textit{Literature Parsing} (L2), \textit{Literature-based Question Answering} (L3), \textit{Literature Review Generation} (L4), and \textit{Scientific Discovery} (L5). HiSciBench contains 8,735 carefully curated instances spanning six major scientific disciplines, including mathematics, physics, chemistry, biology, geography, and astronomy, and supports multimodal inputs including text, equations, figures, and tables, as well as cross-lingual evaluation. Unlike prior benchmarks that assess isolated abilities, HiSciBench provides an integrated, dependency-aware framework that enables detailed diagnosis of model capabilities across different stages of scientific reasoning. Comprehensive evaluations of leading models, including GPT-5, DeepSeek-R1, and several multimodal systems, reveal substantial performance gaps: while models achieve up to 69\% accuracy on basic literacy tasks, performance declines sharply to 25\% on discovery-level challenges. HiSciBench establishes a new standard for evaluating scientific Intelligence and offers actionable insights for developing models that are not only more capable but also more reliable. The benchmark will be publicly released to facilitate future research.

科学智能多模态评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。