一个统一模型同时理解生成跨学科科学数据,支持语言指令预测天气和医学图像。
A unified multimodal understanding and generation model for cross-disciplinary scientific research
- 用统一架构融合多模态科学数据,通过自然语言指令实现跨领域预测
- 生成0.25°分辨率10天全球气象预报,台风路径强度预测优于现有物理模型
- 适用于地球科学与生物医学,适合需要跨学科整合的科研人员
当前科学发现日益依赖跨学科、高维异构数据的整合。尽管人工智能在多个科学领域取得显著进展,但多数模型仍局限于特定领域,缺乏同时理解和生成多模态科学数据的能力,尤其在高维数据上表现不足。而许多重大全球挑战和科学问题本质上具有跨学科性,需多领域协同推进。本文提出FuXi-Uni,一种原生统一的多模态科学理解与生成模型,在单一架构中实现跨科学领域的统一建模。该模型将跨学科科学标记与自然语言标记对齐,并使用科学解码器重建科学标记,从而支持自然语言对话与科学数值预测。实验验证了其在地球科学与生物医学中的性能:在地球系统建模中,仅通过语言指令即可实现全球天气预报、热带气旋(TC)预报编辑与空间降尺度;生成的10天全球预报在0.25°分辨率下优于最先进物理预报系统;在台风路径与强度预测上均超越现有物理模型,生成的高分辨率区域气象场亦优于标准插值基线。在生物医学方面,该模型在多个生物医学视觉问答基准上超越领先多模态大语言模型。通过在原生共享潜在空间中统一异构科学模态,同时保持强领域特定性能,FuXi-Uni推动了更通用的多模态科学模型发展。
原文摘要 · Abstract (English)
Scientific discovery increasingly relies on integrating heterogeneous, high-dimensional data across disciplines nowadays. While AI models have achieved notable success across various scientific domains, they typically remain domain-specific or lack the capability of simultaneously understanding and generating multimodal scientific data, particularly for high-dimensional data. Yet, many pressing global challenges and scientific problems are inherently cross-disciplinary and require coordinated progress across multiple fields. Here, we present FuXi-Uni, a native unified multimodal model for scientific understanding and high-fidelity generation across scientific domains within a single architecture. Specifically, FuXi-Uni aligns cross-disciplinary scientific tokens within natural language tokens and employs science decoder to reconstruct scientific tokens, thereby supporting both natural language conversation and scientific numerical prediction. Empirically, we validate FuXi-Uni in Earth science and Biomedicine. In Earth system modeling, the model supports global weather forecasting, tropical cyclone (TC) forecast editing, and spatial downscaling driven by only language instructions. FuXi-Uni generates 10-day global forecasts at 0.25° resolution that outperform the SOTA physical forecasting system. It shows superior performance for both TC track and intensity prediction relative to the SOTA physical model, and generates high-resolution regional weather fields that surpass standard interpolation baselines. Regarding biomedicine, FuXi-Uni outperforms leading multimodal large language models on multiple biomedical visual question answering benchmarks. By unifying heterogeneous scientific modalities within a native shared latent space while maintaining strong domain-specific performance, FuXi-Uni provides a step forward more general-purpose, multimodal scientific models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。