arXiv:2512.16969cs.AIcs.CL2025-12被引 22

用科学家工作流评估大模型科学通用智能,发现其仍难自主科研。

Probing Scientific General Intelligence of LLMs with Scientist-Aligned Workflows

  • 基于科学探究模型构建四类任务,定义科学通用智能
  • 大模型在深度研究等任务中准确率仅10%-20%,实验结果偏差大
  • 适合研究通用人工智能与科学发现的交叉领域学者

尽管科学人工智能取得进展,但科学通用智能(SGI)——即自主提出、探究并跨领域推理的能力——仍缺乏统一框架。本文基于实践探究模型(PIM:反思、构想、行动、感知),提出四类科学家对齐任务:深度调研、创意生成、干/湿实验与实验推理。SGI-Bench包含1000余个专家标注的跨学科样本,源自《科学》125个重大问题,可系统评估前沿大模型。结果显示:深度调研虽步骤对齐,但精确匹配率仅10%–20%;创意缺乏可行性与细节;干实验代码可执行率高,但执行结果准确率低;湿实验流程序列保真度差;多模态比较推理能力持续薄弱。我们进一步提出测试时强化学习(TTRL),通过检索增强的新颖性奖励优化推理,提升假设新颖性而无需参考答案。本研究的PIM基础定义、以工作流为中心的基准及实证洞察,为真正参与科学发现的AI系统奠定基础。

原文摘要 · Abstract (English)

Despite advances in scientific AI, a coherent framework for Scientific General Intelligence (SGI)-the ability to autonomously conceive, investigate, and reason across scientific domains-remains lacking. We present an operational SGI definition grounded in the Practical Inquiry Model (PIM: Deliberation, Conception, Action, Perception) and operationalize it via four scientist-aligned tasks: deep research, idea generation, dry/wet experiments, and experimental reasoning. SGI-Bench comprises over 1,000 expert-curated, cross-disciplinary samples inspired by Science's 125 Big Questions, enabling systematic evaluation of state-of-the-art LLMs. Results reveal gaps: low exact match (10--20%) in deep research despite step-level alignment; ideas lacking feasibility and detail; high code executability but low execution result accuracy in dry experiments; low sequence fidelity in wet protocols; and persistent multimodal comparative-reasoning challenges. We further introduce Test-Time Reinforcement Learning (TTRL), which optimizes retrieval-augmented novelty rewards at inference, enhancing hypothesis novelty without reference answer. Together, our PIM-grounded definition, workflow-centric benchmark, and empirical insights establish a foundation for AI systems that genuinely participate in scientific discovery.

科学智能大模型评估通用智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。