MEDIC-AD让医学视觉语言模型更懂临床,能精准定位病灶、追踪病情变化并生成可信解释。
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
- 分阶段设计:用异常感知令牌、差异令牌和可解释性模块增强模型临床能力。
- 在纵向数据上实现领先性能,异常检测与分割准确率显著提升。
- 适合临床医生用于患者跟踪与辅助决策,输出结果真实可信。
病灶检测、症状追踪与视觉可解释性是真实医疗影像分析的核心,但现有医学视觉语言模型(VLMs)仍缺乏将广泛知识转化为临床可用输出的能力。为此,我们提出面向临床的MEDIC-AD,通过分阶段框架强化三项能力:首先,可学习的异常感知令牌(<Ano>)引导模型关注异常区域,构建更具区分性的病灶表征;其次,图像间差异令牌(<Diff>)显式编码多期影像的时间变化,使模型能够识别病情恶化、改善或稳定;最后,专门的可解释性阶段训练模型生成与推理一致的热力图,直观呈现病灶相关区域。通过该分阶段设计,MEDIC-AD在异常检测、症状追踪与异常分割任务中持续提升性能,优于多种闭源及医学专用基线模型。在来自真实医院工作流的纵向临床数据上评估显示,MEDIC-AD能提供稳定预测与符合临床逻辑的解释,适用于实际患者监测与决策支持场景。
原文摘要 · Abstract (English)
Lesion detection, symptom tracking, and visual explainability are central to real-world medical image analysis, yet current medical Vision-Language Models (VLMs) still lack mechanisms that translate their broad knowledge into clinically actionable outputs. To bridge this gap, we present MEDIC-AD, a clinically oriented VLM that strengthens these three capabilities through a stage-wise framework. First, learnable anomaly-aware tokens (<Ano>) encourage the model to focus on abnormal regions and build more discriminative lesion centered representations. Second, inter image difference tokens (<Diff>) explicitly encode temporal changes between studies, allowing the model to distinguish worsening, improvement, and stability in disease burden. Finally, a dedicated explainability stage trains the model to generate heatmaps that highlight lesion-related regions, offering clear visual evidence that is consistent with the model's reasoning. Through our staged design, MEDIC-AD steadily boosts performance across anomaly detection, symptom tracking, and anomaly segmentation, achieving state-of-the-art results compared with both closed source and medical-specialized baselines. Evaluations on real longitudinal clinical data collected from real hospital workflows further show that MEDIC-AD delivers stable predictions and clinically faithful explanations in practical patient-monitoring and decision-support workflows
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。