arXiv:2503.01019cs.CVcs.AI2025-03CVPR被引 20

医学多模态模型首次融合图像生成与理解,提升诊疗文本与影像协同能力。

MedUnifier: Unifying Vision-and-Language Pre-training on Medical Data with Vision Generation Task using Discrete Visual Representations

  • 用离散视觉表示替代连续特征,统一医疗图文预训练
  • 在7项任务上达到当前最佳,包括报告生成和零样本分类
  • 适合医疗AI研发者构建通用化诊断辅助系统

尽管视觉-语言预训练(VLP)取得显著进展,但现有方法主要聚焦特征提取与跨模态理解,对视觉内容生成或转换关注不足,限制了模型从文本提示中合成连贯新视觉表征的能力。本文提出针对医疗数据的统一VLP框架MedUnifier,无缝集成文本引导图像生成与多模态学习策略,包括图像-文本对比对齐、匹配及图像引导文本生成。不同于依赖连续视觉表示的传统方法,本方案采用视觉向量量化,不仅促进跨模态理解的一致性学习,还通过离散表示显著提升多模态生成质量。实验在多个基准测试中验证有效性,涵盖单模态(监督微调)、跨模态(图像-文本检索、零样本图像分类)和多模态任务(医疗报告生成、图像合成),均实现最先进性能。MedUnifier为医疗领域语言与视觉任务提供高度可适配工具,推动通用化医疗AI模型的发展。

原文摘要 · Abstract (English)

Despite significant progress in Vision-Language Pre-training (VLP), current approaches predominantly emphasize feature extraction and cross-modal comprehension, with limited attention to generating or transforming visual content. This gap hinders the model's ability to synthesize coherent and novel visual representations from textual prompts, thereby reducing the effectiveness of multi-modal learning. In this work, we propose MedUnifier, a unified VLP framework tailored for medical data. MedUnifier seamlessly integrates text-grounded image generation capabilities with multi-modal learning strategies, including image-text contrastive alignment, image-text matching and image-grounded text generation. Unlike traditional methods that reply on continuous visual representations, our approach employs visual vector quantization, which not only facilitates a more cohesive learning strategy for cross-modal understanding but also enhances multi-modal generation quality by effectively leveraging discrete representations. Our framework's effectiveness is evidenced by the experiments on established benchmarks, including uni-modal tasks (supervised fine-tuning), cross-modal tasks (image-text retrieval and zero-shot image classification), and multi-modal tasks (medical report generation, image synthesis), where it achieves state-of-the-art performance across various tasks. MedUnifier also offers a highly adaptable tool for a wide range of language and vision tasks in healthcare, marking advancement toward the development of a generalizable AI model for medical applications.

医学AI图文生成多模态视觉量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。