arXiv:2607.20557cs.LGcs.AI2026-07

一个统一的科学多模态模型,能跨领域理解与生成生物、化学、气象等数据。

Monkey King Bang: A Unified Scientific Multimodal Foundation Model

论文配图:Monkey King Bang: A Unified Scientific Multimodal Foundation Model
图 1 · 摘自论文原文
  • 共享变压器主干+专用编码器/解码器,适配六大学科数据输入输出
  • 在生物分子和气象预测任务中达到领先水平,生成结果符合真实模态特性
  • 适合科研人员做跨学科智能分析,也适用于多模态科学生成任务

科学发现正从单一学科转向多领域协同推理,人工智能在科学领域的应用也面临类似转变。现有系统或局限于特定领域,或仅通过文本分词与提示接口整合科学数据,难以处理多样化的科学输入,无法生成原生模态输出,也无法支持跨领域的联合理解、推理与生成。我们提出MKB,一个统一的科学多模态模型,具备理解与生成能力,基于共享Transformer主干,搭配针对不同模态设计的编码器、适配器与解码器。MKB覆盖六个科学分支:DNA、RNA、蛋白质、小分子、地球科学与医学影像,支持原生输出如生物序列、分子字符串、气象场与分割掩码。训练采用两阶段课程:第一阶段对齐各模态组件与冻结主干;第二阶段使用混合科学与通用语料,整合语言主干。实验表明,MKB在生物与分子基准测试中表现优异,生成的天气预报、生物序列与医学图像分割结果具有高保真度,且基本保留其Qwen3-VL主干的通用能力。结果验证了该范式的可行性,表明共享主干结合模态定制组件可为未来跨领域科学多模态探索提供有力基础。模型与代码已开源:https://github.com/Shanghai-Academy-of-AI-For-Science/MKB 与 https://huggingface.co/sais-org/MKB。

原文摘要 · Abstract (English)

Scientific discovery is increasingly shifting from isolated disciplines to multi-domain reasoning, and AI for science faces a similar transition. Existing systems are either specialised for individual domains or unify scientific data mainly through text tokenisation and prompt-based interfaces, limiting their ability to handle diverse scientific inputs, produce modality-native outputs, and support joint understanding, reasoning, and generation across scientific domains. We introduce MKB, a unified scientific multimodal model for both understanding and generation, built around a shared Transformer backbone and modality-tailored encoders, adapters, and decoders. MKB covers six scientific branches, including DNA, RNA, proteins, small molecules, earth science, and medical images, and supports native outputs such as biological sequences, molecular strings, meteorological fields, and segmentation masks. Training follows a two-stage modality-then-language curriculum: Stage 1 aligns modality-specific components with the frozen backbone, and Stage 2 consolidates them with the language backbone using mixed scientific and general corpora. Experiments show that MKB achieves competitive scientific understanding across biological and molecular benchmarks, produces high-fidelity native outputs for weather forecasting, biological generation, and medical-image segmentation, and largely retains the general capabilities of its Qwen3-VL backbone. These results demonstrate the feasibility of the proposed paradigm, suggesting that shared-backbone models with modality-tailored components can provide a promising foundation for future cross-domain scientific multimodal exploration. The model and code are publicly available at https://github.com/Shanghai-Academy-of-AI-For-Science/MKB and https://huggingface.co/sais-org/MKB.

多模态科学计算生成模型跨领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。