arXiv:2511.20732cs.MMcs.CV2025-11中稿 · 32nd International…被引 1

针对医学视觉语言模型的持续学习遗忘问题,提出按提示分组保护关键参数的新方法。

Prompt-Aware Adaptive Elastic Weight Consolidation for Continual Learning in Medical Vision-Language Models

  • 根据图像-描述、空间引导和医学语义功能分组参数,实现精准保护。
  • 在5个数据集上减少遗忘17.58%,提升肺部病变定位与息肉分割性能。
  • 适合需长期更新但不能丢弃历史诊断能力的临床AI系统使用。

医疗AI系统在临床部署中面临灾难性遗忘问题,即学习新影像协议时会丢失原有诊断能力。这一挑战尤其突出在需要跨模态对齐医学图像与临床术语的视觉语言模型中。本文提出提示感知自适应弹性权重巩固(PA-EWC),通过提示引导的参数专业化缓解遗忘。该方法基于视觉-描述、空间引导和医学语义三类功能对模型参数进行系统分类,从而针对性保护关键知识,同时支持对新临床需求的适应。PA-EWC结合自适应费雪信息计算与梯度稳定性分析,并引入基于医学术语密度的加权复杂度度量。我们在五个医学影像数据集(Kvasir-SEG, ISIC 2018, CheXlocalize, BUSI, CAMUS)上评估,涵盖内窥镜、皮肤镜、胸片和超声等多种模态。实验表明,相比基线方法,PA-EWC将灾难性遗忘降低17.58%,在胸部X光病理定位任务上提升4.30%,在息肉分割任务上提升6.06%。

原文摘要 · Abstract (English)

Medical AI systems face catastrophic forgetting when deployed in clinical settings, where models must learn new imaging protocols while retaining prior diagnostic capabilities. This challenge is particularly acute for medical vision-language models that must preserve complex cross-modal alignments between medical images and clinical terminology across diverse imaging modalities. We introduce Prompt- Aware Adaptive Elastic Weight Consolidation (PA-EWC), a novel continual learning approach that addresses catastrophic forgetting through prompt-guided parameter specialization. Our method systematically categorizes model parameters based on their functional roles in processing visual-descriptive, spatial-guided, and medical-semantic information, enabling targeted protection of critical knowledge while allowing adaptation to new clinical requirements. PA-EWC incorporates adaptive Fisher Information computation with gradient stability analysis and develops weighted complexity metrics based on medical terminology density. We evaluate our approach across five medical imaging datasets (Kvasir-SEG, ISIC 2018, CheXlocalize, BUSI, CAMUS) representing diverse modalities including endoscopy, dermoscopy, radiography, and ultrasound. Experimental results demonstrate that PA-EWC reduces catastrophic forgetting by up to 17.58% compared to baseline methods, with performance improvements of 4.30% on chest X-ray pathology localization and 6.06% on polyp segmentation.

持续学习医学视觉语言参数保护灾难性遗忘

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。