arXiv:2603.25040cs.LGcs.CL2026-03被引 16

千亿参数科学多模态模型,通用与专业能力兼备。

Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale

  • 千亿参数规模,融合通用与科学领域知识
  • 掌握100+科学任务,覆盖化学、材料等关键领域
  • 支持高效强化学习训练,开源模型中顶尖性能

我们推出Intern-S1-Pro,首个达到万亿参数规模的科学多模态基础模型。该模型在前所未有的规模下,显著提升通用与科学领域的综合能力。不仅增强推理与图文理解能力,还具备先进的智能体功能。其科学专长扩展至100余项关键科学任务,涵盖化学、材料、生命科学和地球科学等领域。通过XTuner和LMDeploy的稳定基础设施支持,实现万亿参数级别的高效强化学习训练,并确保训练与推理间的严格精度一致性。通过整合这些进步,Intern-S1-Pro进一步强化通用与专业智能的融合,作为可定制的通用专家,在通用能力上跻身开源模型前列,同时在科学任务深度上超越多数闭源模型。

原文摘要 · Abstract (English)

We introduce Intern-S1-Pro, the first one-trillion-parameter scientific multimodal foundation model. Scaling to this unprecedented size, the model delivers a comprehensive enhancement across both general and scientific domains. Beyond stronger reasoning and image-text understanding capabilities, its intelligence is augmented with advanced agent capabilities. Simultaneously, its scientific expertise has been vastly expanded to master over 100 specialized tasks across critical science fields, including chemistry, materials, life sciences, and earth sciences. Achieving this massive scale is made possible by the robust infrastructure support of XTuner and LMDeploy, which facilitates highly efficient Reinforcement Learning (RL) training at the 1-trillion parameter level while ensuring strict precision consistency between training and inference. By seamlessly integrating these advancements, Intern-S1-Pro further fortifies the fusion of general and specialized intelligence, working as a Specializable Generalist, demonstrating its position in the top tier of open-source models for general capabilities, while outperforming proprietary models in the depth of specialized scientific tasks.

多模态科学智能大模型万亿参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。