arXiv:2511.00916cs.CV2025-11被引 6

构建统一医学视觉理解框架,支持多模态医疗数据的通用推理。

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs

  • 从数据出发,融合自然与医学领域长文本提升预训练效果
  • 覆盖超声、皮肤镜等罕见模态,增强对视频和3D影像的理解能力
  • 开源模型支持临床研究,推动可复现的医疗AI发展

多模态大语言模型在通用场景中表现优异,但医学数据因模态多样(2D图像、3D体数据、时序视频)且格式不一,导致统一建模困难。本文提出Fleming-VL,一个面向异构医学模态的端到端统一框架。通过三方面策略:(1)融合自然与医学领域的长上下文数据扩大预训练规模;(2)引入罕见医学数据,涵盖整体视频分析及超声、皮肤镜等低频2D模态;(3)扩展评估体系,纳入3D体数据与视频理解基准。采用监督微调(SFT)与组相对策略优化(GRPO)训练多尺度模型。实验表明,Fleming-VL在医学VQA、视频问答与3D医学图像理解等多个基准上达到当前最优性能。模型已公开发布,促进医疗AI透明、可复现、可审计的发展。

原文摘要 · Abstract (English)

Multimodal Large Language Models (MLLMs) have demonstrated remarkable effectiveness in various general-domain scenarios, such as visual question answering and image captioning. Recently, researchers have increasingly focused on empowering MLLMs with medical conversational abilities, which hold significant promise for clinical applications. However, medical data presents unique challenges due to its heterogeneous nature -- encompassing diverse modalities including 2D images, 3D volumetric scans, and temporal video sequences. The substantial domain gap and data format inconsistencies across these modalities have hindered the development of unified medical MLLMs. To address these challenges, we propose Fleming-VL, a unified end-to-end framework for comprehensive medical visual understanding across heterogeneous modalities. Fleming-VL tackles this problem from a data-centric perspective through three key strategies: (1) scaling up pretraining by integrating long-context data from both natural and medical-specific domains; (2) complementing fine-tuning with rare medical data, including holistic video analysis and underrepresented 2D modalities such as ultrasound and dermoscopy images; (3) extending existing evaluation frameworks to incorporate 3D volumetric and video understanding benchmarks. Through supervised fine-tuning (SFT) and group relative policy optimization (GRPO), we develop Fleming-VL in multiple model scales. Extensive experiments demonstrate that Fleming-VL achieves state-of-the-art performance across multiple benchmarks, including medical VQA, video QA, and 3D medical image understanding. We publicly release Fleming-VL to promote transparent, reproducible, and auditable progress in medical AI.

医学视觉多模态大模型视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。