用统一语言模型生成科学内容,一模型通吃多领域任务
Speaking the Language of Science: Toward a General-Purpose Generative Foundation Model for the Natural Sciences

- 将科学对象与空间关系转为通用词元序列,用纯序列建模复杂结构
- 在10余项科学任务中表现超越专用模型,大模型效果更优
- 适合想用大模型做科研的学者,推动科学与大模型深度融合
本文提出LOGOS(Language Of Generative Objects in Science),一种统一自然科学研究中异构任务的生成式语言模型。它基于共享科学语法,将多样化的科学对象及其空间相互作用编码为单一词汇表下的词元序列。通过将空间接触与约束模式表示为离散词元,模型以纯序列方式捕捉复杂结构交互,无需显式坐标或几何神经网络。该统一表示使多种下游任务可一致地表述为同一语法空间中的下一步词元预测,实现持续跨领域预训练与下游目标间的强对齐。在多样化任务中,LOGOS始终达到或超过领域专用基线表现,初步验证了“一模型通吃”在自然科学中的可行性。我们训练了1B、3B和8B参数规模的LOGOS模型,发现模型规模与性能呈稳定正相关。这表明,人工智能赋能科学(AI4S)的未来可能不在于构建脱离大语言模型(LLMs)的独立技术栈,而在于通过共享架构、共享训练范式与共享推理基础设施,深度对齐科学基础模型与大语言模型,使LLMs真正成为进入AI4S的新入口。模型权重及相关资源已公开。
原文摘要 · Abstract (English)
In this report, we present LOGOS (Language Of Generative Objects in Science), a scientific generative language model that unifies heterogeneous tasks across the natural sciences within a single autoregressive framework based on a shared scientific grammar. It encodes diverse scientific objects and their spatial interactions as token sequences over a common vocabulary. By representing spatial contact and constraint patterns as discrete tokens, the model captures complex structural interactions in a purely sequential manner, without relying on explicit coordinates or geometric neural networks. This unified representation enables a wide range of downstream tasks to be formulated consistently as next-token prediction in the same grammar space, creating strong alignment between continued multi-domain pre-training and downstream objectives. Across diverse tasks, LOGOS consistently matches or outperforms domain-specific baselines, providing preliminary evidence for the feasibility of "one model fits all" in the natural sciences. We train LOGOS models at different scales (1B, 3B, and 8B parameters) and find a consistent positive correlation between model size and performance. This suggests that the future of AI for Science (AI4S) may not lie in building an independent technical stack that is separated from large language models (LLMs). Instead, it may depend on deeply aligning scientific foundation models with LLMs through shared architectures, shared training paradigms, and shared inference infrastructure, so that LLMs can truly become a new entry point for AI4S. We release the model weights and associated resources to facilitate further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。