提出四角色框架,解析大模型在科研中的定位与边界
Evolving Roles of LLMs in Scientific Innovation: Assistant, Collaborator, Scientist, and Evaluator
- 按自主性、认知功能、创新维度构建四角色框架
- 助手类成熟但开放应用不可靠,科学家类存在安全瓶颈
- 适合关注AI科研工具设计与治理的研究者
大语言模型正广泛应用于科学研宄与发现,支持文献检索与综述、假设生成、自主实验及研究评估等任务。现有综述常混淆科研与科学发现,仅按领域、任务或自主性组织系统。本文提出四角色框架:助理、合作者、科学家与评估者,融合自主性、认知功能与科学创新三维度,区分研究支持与前沿探索。综述各角色代表性方法、基准与评估实践,分析其能力、局限与人工监管需求。文献显示,助理系统在检索与综述上较成熟,但在开放式任务中不可靠;合作者扩展假设空间,但面临新颖性权衡难题;科学家系统逐步自动化研究流程,却受限于可靠性和安全性;评估系统辅助评审与验证,但在新颖性判断上仍薄弱。我们认为,科学AI进步不仅依赖模型能力,更需评估机制、监督体系、问责制度与机构整合。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly used in scientific research and discovery, supporting tasks ranging from literature retrieval and synthesis to hypothesis generation, autonomous experimentation, and research evaluation. Existing surveys often conflate scientific research with scientific discovery and typically organize systems by domain, task, or autonomy level alone. In this survey, we propose a four-role framework for understanding LLMs in scientific innovation: Assistant, Collaborator, Scientist, and Evaluator. The framework integrates three complementary dimensions: autonomy level, cognitive function, and scientific innovation, to distinguish research-oriented support from frontier-oriented discovery. We review representative methods, benchmarks, and evaluation practices for each role, examining their capabilities, limitations, and human oversight requirements. Across the literature, Assistant systems are comparatively mature in retrieval and synthesis but remain unreliable in open-ended applications; Collaborator systems expand the space of candidate hypotheses yet struggle with novelty-grounding trade-offs; Scientist systems increasingly automate research workflows but face reliability and safety bottlenecks; and Evaluator systems support review and verification while remaining weak in novelty assessment. We argue that progress in AI for science depends not only on model capability, but also on evaluation, oversight, accountability, and institutional integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。