系统梳理科学大模型的数据基础与智能演化,揭示其发展核心是数据与模型的协同进化。
A Survey of Scientific Large Language Models: From Data Foundations to Agent Frontiers
- 以数据为中心重构科学大模型发展路径,提出多模态跨尺度知识层级框架。
- 分析270+预/后训练数据集与190+评测基准,发现科学数据具异构性与不确定性特征。
- 展望闭环智能体系统,支持自主实验与动态知识更新,适合科研自动化研究者。
科学大语言模型(Sci-LLMs)正在重塑科学知识的表征、整合与应用方式,但其进展受制于科学数据的复杂性。本文从数据出发,将Sci-LLMs的发展重构为模型与底层数据基底的协同演化过程。提出统一的科学数据分类体系与分层科学知识模型,强调科学语料在多模态、跨尺度、领域特异性方面的挑战,区别于通用自然语言处理数据集。系统综述了涵盖多个科学领域的通用与专用模型,并分析超过270个预/后训练数据集,揭示其对异构、多尺度、含不确定性的数据需求,要求模型具备领域不变性表示与跨模态推理能力。在评估方面,考察190余个基准数据集,呈现从静态测试向过程与发现导向评估的转变,伴随先进评估协议的发展。这些数据驱动分析揭示了科学数据构建中的持续问题,并讨论半自动标注流程与专家验证等新兴解决方案。最后,提出向闭环系统范式演进:基于Sci-LLMs的自主智能体可主动实验、验证并贡献于动态演化的知识库。本工作为构建可信、持续进化的AI系统提供路线图,使其真正成为加速科学发现的伙伴。
原文摘要 · Abstract (English)
Scientific Large Language Models (Sci-LLMs) are transforming how knowledge is represented, integrated, and applied in scientific research, yet their progress is shaped by the complex nature of scientific data. This survey presents a comprehensive, data-centric synthesis that reframes the development of Sci-LLMs as a co-evolution between models and their underlying data substrate. We formulate a unified taxonomy of scientific data and a hierarchical model of scientific knowledge, emphasizing the multimodal, cross-scale, and domain-specific challenges that differentiate scientific corpora from general natural language processing datasets. We systematically review recent Sci-LLMs, from general-purpose foundations to specialized models across diverse scientific disciplines, alongside an extensive analysis of over 270 pre-/post-training datasets, showing why Sci-LLMs pose distinct demands -- heterogeneous, multi-scale, uncertainty-laden corpora that require representations preserving domain invariance and enabling cross-modal reasoning. On evaluation, we examine over 190 benchmark datasets and trace a shift from static exams toward process- and discovery-oriented assessments with advanced evaluation protocols. These data-centric analyses highlight persistent issues in scientific data development and discuss emerging solutions involving semi-automated annotation pipelines and expert validation. Finally, we outline a paradigm shift toward closed-loop systems where autonomous agents based on Sci-LLMs actively experiment, validate, and contribute to a living, evolving knowledge base. Collectively, this work provides a roadmap for building trustworthy, continually evolving artificial intelligence (AI) systems that function as a true partner in accelerating scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。