arXiv:2606.24441cs.CV2026-06被引 1

统一模型实现科学图像理解、生成与编辑,支持多轮编辑与专业可视化。

S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing

论文配图:S1-Omni-Image: A Unified Model for Scientific Image Understanding, Generation, and Editing
图 1 · 摘自论文原文
  • 先推理再生成,通过任务追踪引导图像创作
  • 在4个编辑基准上达顶尖性能,生成符合科学逻辑的图表与图像
  • 适合科研人员、医学影像与工程可视化场景使用

我们提出S1-Omni-Image,一个开源权重的统一多模态模型,用于科学图像的理解、生成与编辑。不同于通用图像生成模型,科学图像任务不仅要求高保真合成,还需精准理解科学语义、结构关系、领域知识和任务意图。为此,S1-Omni-Image基于科学多模态推理主干S1-VL-32B,采用统一的‘先思考后生成’范式,将理解能力与图像生成模块耦合。给定用户指令后,模型首先生成任务导向的推理轨迹、文本答案及任务专属标记,其隐状态被注入生成模块以条件化图像生成或编辑。该模型支持科学图像理解、生成与编辑的统一框架。生成方面聚焦科学插图与文本渲染,包括逻辑图、关系对比图、数据图表和真实科学可视化;编辑方面将分割等专用视觉任务转化为原生图像编辑问题,支持多轮插图编辑、医疗与地理图像分割、医学图像翻译和科学图像超分辨率。我们构建了包含31.4万样本的SciGenEdit训练数据集,并发布模型权重、推理代码及SciGenEdit-10K。实验表明,S1-Omni-Image显著提升科学图像生成与编辑性能,同时保持来自S1-VL-32B的科学理解能力。其在GenExam和TechImage-Bench上超越开源模型,在四个编辑基准(MSD、cigRockSEM、SynthRAD2025、IXI)上达到顶尖水平,并在科学图像理解评估中表现稳定。

原文摘要 · Abstract (English)

We present S1-Omni-Image, an open-weight unified multimodal model for scientific image understanding, generation, and editing. Unlike general-purpose image generation models, scientific image tasks require not only high-fidelity synthesis, but also robust understanding of scientific semantics, structural relations, domain knowledge, and task intent. To this end, S1-Omni-Image builds on the scientific multimodal reasoning backbone S1-VL-32B and couples its understanding capability with an image generation module under a unified think-before-generate paradigm. Given a user instruction, the model first produces a task-oriented reasoning trace, a textual answer, and a task special token; their hidden states are then injected into the generation module to condition image generation or editing. S1-Omni-Image supports scientific image understanding, generation, and editing in a unified framework. For generation, it focuses on scientific illustrations and text rendering, including logical diagrams, relational comparisons, data charts, and realistic scientific visualizations. For editing, it casts segmentation and other domain-specific vision tasks as native image editing problems, enabling multi-turn illustration editing, medical and geographic image segmentation, medical image translation, and scientific image super-resolution. We construct SciGenEdit, a 314K-sample training dataset, and release the model weights, inference code, and SciGenEdit-10K. Experiments show that S1-Omni-Image substantially improves scientific image generation and editing while preserving the scientific image understanding capability inherited from S1-VL-32B. It outperforms open-source models on GenExam and TechImage-Bench, achieves state-of-the-art results on four editing benchmarks including MSD, cigRockSEM, SynthRAD2025, and IXI, and maintains stable performance on scientific image understanding evaluations.

科学图像统一模型图像生成图像编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。