arXiv:2511.22256cs.CV2025-11被引 2

打造首个统一超声图像理解与临床推理的通用模型

UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation

  • 用动态卷积掩码解码器实现像素级结构理解
  • 在16个解剖区域上达顶尖诊断与分割性能
  • 适合医疗AI研发与临床辅助系统开发者

尽管医学基础模型进展显著,超声领域仍缺乏能融合低层结构感知(如分割、定位)与高层综合解读(如诊断、推理)的完整解决方案。为此,我们提出UMind-VL,一个统一的基础模型,旨在协同实现像素级结构理解与复杂临床推理。我们首先构建了包含120万张超声图像-文本对的大型多模态数据集UMind-DS,覆盖16个解剖区域,包含像素级标注和临床医生验证的推理依据。架构上,UMind-VL引入轻量级动态卷积掩码解码器,通过基于大语言模型输出的动态核生成掩码,结合任务特定标记,将分割、检测、几何测量和诊断任务统一于单一框架。大量评估表明,UMind-VL显著优于现有通用多模态模型,并在分割、检测、关键点定位和诊断推理基准上达到或超过最先进专用模型的性能,同时保持强泛化能力。

原文摘要 · Abstract (English)

Despite significant strides in medical foundation models, the ultrasound domain lacks a comprehensive solution capable of bridging low-level Ultrasound Grounded Perception (e.g., segmentation, localization) and high-level Ultrasound Comprehensive Interpretation (e.g., diagnosis, reasoning). To bridge this gap, we propose UMind-VL, a unified foundation model designed to synergize pixel-level structural understanding with complex clinical reasoning. We first introduce UMind-DS, a large-scale multimodal dataset comprising 1.2 million ultrasound image-text pairs across 16 anatomical regions, enriching standard data with pixel-level annotations and clinician-validated rationales. Architecturally, UMind-VL incorporates a lightweight Dynamic Convolutional Mask Decoder that generates masks via dynamic kernels conditioned on LLM outputs. This design, combined with task-specific tokens, unifies segmentation, detection, geometric measurement, and diagnosis tasks within a single framework. Extensive evaluations demonstrate that UMind-VL significantly outperforms existing generalist multimodal models and achieves performance on par with, or superior to, state-of-the-art specialist models across segmentation, detection, keypoint localization, and diagnostic reasoning benchmarks, while maintaining strong generalization ability. We demonstrate the capability of UMind-VL in Figure 1.

超声分析多模态临床推理通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。