arXiv:2510.15710cs.CV2025-10被引 13

首个统一医学多模态理解与生成的模型,支持图像与文本协同处理。

UniMedVL: Unifying Medical Multimodal Understanding and Generation through Observation-Knowledge-Analysis

  • 通过渐进式训练让理解与生成能力相互增强,实现单模型统一。
  • 在5个医学理解基准上表现优秀,支持虚拟染色等7种交叉生成任务。
  • 适合需要图像与文本联合处理的临床AI系统研发者使用。

医疗工作流程通常结合图像阅读与视觉/文本输出生成,使多模态理解与生成成为医学AI的核心。现有系统大多采用孤立模型分别处理,错失统一架构带来的共享知识优势。为此,我们提出UniMedVL,首个在单一模型中无缝集成多模态理解与生成能力的医学模型,无需权重切换。通过定制的渐进式训练流程,理解与生成相互促进。为有效训练,我们构建了包含超过560万样本、覆盖8种医学影像模态的UniMedVL-5M数据集,专为统一医学理解与生成任务设计。实验表明,UniMedVL在5个医学理解基准上达到有竞争力的表现。关键的是,该模型原生支持多种交错生成任务,如虚拟染色、超分辨率、跨模态合成,对复杂医疗流程至关重要。代码与数据集已公开。

原文摘要 · Abstract (English)

Medical workflows routinely combine reading images with producing visual and textual outputs, making both image understanding and generation central to medical AI. Most existing systems, however, address these abilities in isolated models, losing the shared knowledge that a unified architecture could exploit. To bridge this gap, we present UniMedVL, the first unified medical model that seamlessly integrates multimodal understanding and generation capabilities within a single model without switching weights. We achieve this via a tailored progressive training pipeline where understanding and generation mutually reinforce each other. To effectively train UniMedVL, we curate UniMedVL-5M, the first large-scale medical dataset comprising over 5.6M instances across 8 medical imaging modalities, tailored for multimodal input-output tasks in unified medical understanding and generation. Experimental results demonstrate that UniMedVL achieves competitive performance on five medical understanding benchmarks. Crucially, UniMedVL natively supports diverse interleaved generation tasks, e.g., virtual staining, super-resolution, cross-modal synthesis, essential for complex medical workflows. Our code and dataset are publicly available.

医学多模态生成模型统一架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。