arXiv:2607.09135cs.CV2026-07

融合通用与专精模型,提升医学影像诊断的全面性与准确性。

Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy

论文配图:Super-Generalist: Towards Comprehensive and Accurate Medical Image Understanding via Generalist-Specialist Synergy
图 1 · 摘自论文原文
  • 通过多专家分割先验增强视觉语言对齐,实现细粒度解剖与病灶感知。
  • 在胸部和腹部CT多个基准上达到顶尖性能,肿瘤诊断优于专用模型。
  • 具备强病灶定位能力,可泛化至无监督类别病灶,适合临床部署。

医学影像需要全面且精确的解读以支持多种临床疾病的诊断。近期的视觉-语言通用模型虽具备广泛的任务覆盖和出色的零样本能力,但往往缺乏细粒度的解剖结构与病灶意识,影响诊断可靠性与空间可解释性。相比之下,监督训练的专用模型在特定任务上表现强劲,但跨疾病与解剖结构的泛化能力有限。本文提出SuG(Super-Generalist)框架,将通用视觉-语言学习与专用目标相结合,实现广谱泛化与专用级诊断能力。SuG通过引入多个分割专家的空间先验(包括解剖、特定病灶及非特定病灶分割器),增强视觉-语言对齐,捕捉训练中未标注的病灶区域。为提升病灶定位能力,利用病灶掩码作为空间先验,校准文本条件下的视觉注意力,使疾病语义聚焦于临床相关区域。我们在多个胸部和腹部CT基准(包括CT-RATE、Merlin、MedVL-CT69K及若干自建肿瘤数据集)上评估了SuG,结果表明其在广泛疾病诊断任务中均达领先水平,并在多个关键肿瘤诊断基准上超越专用模型。此外,SuG展现出优异的病灶定位能力,能有效泛化至未受监督的病灶类型。

原文摘要 · Abstract (English)

Medical images require comprehensive and accurate interpretation to support the diagnosis of diverse clincial conditions. Recent vision-language generalist models offer broad task coverage and promising zero-shot capabilities, yet often lack fine-grained anatomical and lesion awareness for reliable diagnosis and spatial interpretability. In contrast, supervised specialist models achieve strong performance on specific tasks but typically lack generalization across diseases and anatomies. In this work, we present SuG, a Super-Generalist framework that unifies generalist vision-language learning with specialist objectives, enabling both broad generalization and specialist-level diagnostic capability. We perform specialist-enhanced vision-language alignment in SuG by incorporating spatial priors from multiple segmentation experts, including anatomy, class-specific lesion and class-agnostic lesion segmentors that captures lesions beyond anatomies annotated during training. To improve lesion grounding capability, we leverage lesion masks as spatial priors to calibrate text-conditioned visual attention, encouraging disease-related semantics to focus on clinically relevant regions. We evaluate SuG on extensive chest and abdominal CT benchmarks, including CT-RATE, Merlin, MedVL-CT69K, and several in-house tumor datasets. SuG achieves state-of-the-art performance across a wide range of disease diagnosis tasks and surpasses specialist models on several critical tumor diagnosis benchmarks. Furthermore, SuG demonstrates strong lesion grounding capability, including robust generalization to lesion types lacking class-specific supervision.

医学影像视觉语言病灶定位通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。