arXiv:2503.13503cs.LGcs.CL2025-03KDD被引 26

构建科学领域AI评估框架,覆盖数据质量与模型能力双维度。

SciHorizon: Benchmarking AI-for-Science Readiness from Scientific Data to Large Language Models

  • 从数据质量、FAIR性等15个子维度评估科研数据可用性
  • 评测超50个主流大模型在5大学科的综合能力表现
  • 提供跨学科数据推荐清单,助力科研AI落地

近年来,人工智能(尤其是大语言模型)的快速发展推动了科学发现范式的变革,使科学智能(AI4Science)成为快速演进的前沿领域。然而,现有缺乏从数据到模型的全链条评估体系。本文提出SciHorizon,一个涵盖科学数据与大模型能力的综合性评估框架。首先,建立可推广的数据评估框架,包含质量、可发现性、可解释性与合规性四个核心维度,共15个子维度;基于2018至2023年发表于同行评审期刊的文献资源,为地球、生命与材料科学构建了首批可直接用于AI训练的高质量数据集推荐列表。同时,针对大模型能力,设计涵盖知识、理解、推理、多模态与价值观的16个评估维度,覆盖数学、物理、化学、生命科学及地球与空间科学。基于自建基准数据集,对超过50个开源与闭源大模型进行了系统评测,所有结果均公开可查,网址:www.scihorizon.cn/en。

原文摘要 · Abstract (English)

In recent years, the rapid advancement of Artificial Intelligence (AI) technologies, particularly Large Language Models (LLMs), has revolutionized the paradigm of scientific discovery, establishing AI-for-Science (AI4Science) as a dynamic and evolving field. However, there is still a lack of an effective framework for the overall assessment of AI4Science, particularly from a holistic perspective on data quality and model capability. Therefore, in this study, we propose SciHorizon, a comprehensive assessment framework designed to benchmark the readiness of AI4Science from both scientific data and LLM perspectives. First, we introduce a generalizable framework for assessing AI-ready scientific data, encompassing four key dimensions: Quality, FAIRness, Explainability, and Compliance-which are subdivided into 15 sub-dimensions. Drawing on data resource papers published between 2018 and 2023 in peer-reviewed journals, we present recommendation lists of AI-ready datasets for Earth, Life, and Materials Sciences, making a novel and original contribution to the field. Concurrently, to assess the capabilities of LLMs across multiple scientific disciplines, we establish 16 assessment dimensions based on five core indicators Knowledge, Understanding, Reasoning, Multimodality, and Values spanning Mathematics, Physics, Chemistry, Life Sciences, and Earth and Space Sciences. Using the developed benchmark datasets, we have conducted a comprehensive evaluation of over 50 representative open-source and closed source LLMs. All the results are publicly available and can be accessed online at www.scihorizon.cn/en.

科学智能大模型评测数据质量基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。