一个统一模型,让AI能理解、预测并生成科学内容。
S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

- 用统一空间表示科学数据与语言指令,支持多模态输入
- 在百万级样本上训练,在60多个科学基准上超越主流模型
- 适合科研人员和开发者用于跨领域科学任务
我们提出S1-Omni,一个统一的多模态科学推理模型,用于科学理解、预测与生成。尽管人工智能赋能科学(AI4S)已通过领域专用模型、工具增强的大模型及科学语言模型取得进展,但模型能力仍高度分散,难以联合建模异构数据、科学规律与专家知识。S1-Omni通过整合三类核心组件解决此问题:统一的科学数据表征、自然世界知识对齐、以及面向特定任务的解码机制。首先,将自然语言指令与科学对象(如CIF、SMILES、蛋白质序列、光谱、科学图像)映射至共享表示空间;其次,将科学定律与专家知识融入数据构建与训练过程,使模型能基于科学证据推理;第三,采用任务特异性解码,支持属性预测、谱图到分子生成、蛋白质位点与结构预测、科学图像生成与编辑等广泛应用。S1-Omni在覆盖200项科学任务的S1-Omni-Corpus上训练,包含数百万推理样本,并在60余个科学基准上评估。其性能优于GPT-5.5与Gemini-3.1-Pro,且在多个基准上达到或超过领域专用模型水平。总体而言,S1-Omni为统一科学建模提供了可行路径。
原文摘要 · Abstract (English)
We present S1-Omni, a unified multimodal reasoning model for scientific understanding, prediction, and generation. AI for Science (AI4S) has advanced significantly through domain-specific models, tool-augmented LLMs, and scientific language models. However, model capabilities remain highly fragmented, limiting the joint modeling of heterogeneous data, scientific laws, and expert knowledge. S1-Omni addresses this gap by consolidating these capabilities into a single, coherent scientific reasoning model. The architecture of S1-Omni is built upon three core components: unified representation of scientific data, natural-world knowledge alignment, and decoding for domain-specific tasks. First, S1-Omni maps natural-language instructions and scientific objects, including CIF, SMILES, protein sequences, spectra, and scientific images, into a shared representation space. Second, it incorporates scientific laws and expert knowledge into data construction and training, enabling the model to reason from scientific evidence. Third, it performs task-specific decoding to support a broad range of applications, including property prediction, spectrum-to-molecular generation, protein site and structure prediction, and scientific image generation and editing. S1-Omni is trained on S1-Omni-Corpus, which covers 200 scientific tasks and contains millions of reasoning samples, and is evaluated on over 60 scientific benchmarks. It outperforms GPT-5.5 and Gemini-3.1-Pro on most benchmarks and matches or surpasses domain-specific models on several benchmarks. Overall, S1-Omni provides a practical path toward unified scientific modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。