arXiv:2504.14692cs.CL2025-04被引 20

统一处理医学多模态数据,提升模型效率与性能。

OmniV-Med: Scaling Medical Vision-Language Model for Universal Visual Understanding

  • 用统一架构处理2D/3D图像和视频,避免分模态编码。
  • 在7个基准上达顶尖水平,1.5B小模型仅需8张3090卡训练。
  • 通过剪枝减少60%视觉令牌,适合医疗视频长序列推理。

医疗视觉语言模型(Med-VLMs)的实际部署需要无缝整合文本与多种视觉模态(包括2D/3D图像和视频),但现有模型通常对不同模态使用独立编码器。为此,我们提出OmniV-Med,一个统一的多模态医学理解框架。技术贡献有三:首先,构建OmniV-Med-Instruct数据集,包含25.2万条指令样本,覆盖14种医学影像模态和11项临床任务;其次,设计旋转变位置自适应编码器,统一处理多分辨率2D/3D图像与视频;第三,引入医学感知的标记剪枝机制,利用体数据(如连续CT切片)和医学视频中的时空冗余,有效减少60%视觉标记而无性能损失。实证表明,OmniV-Med-7B在7个涵盖2D/3D医学影像与视频理解的基准上达到当前最优性能。其轻量版OmniV-Med-1.5B表现相当,训练仅需8张RTX3090 GPU,并支持高效长视频推理。数据、代码与模型将公开。

原文摘要 · Abstract (English)

The practical deployment of medical vision-language models (Med-VLMs) necessitates seamless integration of textual data with diverse visual modalities, including 2D/3D images and videos, yet existing models typically employ separate encoders for different modalities. To address this limitation, we present OmniV-Med, a unified framework for multimodal medical understanding. Our technical contributions are threefold: First, we construct OmniV-Med-Instruct, a comprehensive multimodal medical dataset containing 252K instructional samples spanning 14 medical image modalities and 11 clinical tasks. Second, we devise a rotary position-adaptive encoder that processes multi-resolution 2D/3D images and videos within a unified architecture, diverging from conventional modality-specific encoders. Third, we introduce a medical-aware token pruning mechanism that exploits spatial-temporal redundancy in volumetric data (e.g., consecutive CT slices) and medical videos, effectively reducing 60\% of visual tokens without performance degradation. Empirical evaluations demonstrate that OmniV-Med-7B achieves state-of-the-art performance on 7 benchmarks spanning 2D/3D medical imaging and video understanding tasks. Notably, our lightweight variant (OmniV-Med-1.5B) attains comparable performance while requiring only 8 RTX3090 GPUs for training and supporting efficient long-video inference. Data, code and model will be released.

医学视觉多模态模型压缩视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。