用大模型生成视觉概念,提升医疗图像持续学习效果
Augmenting Continual Learning of Diseases with LLM-Generated Visual Concepts
- 用大模型生成疾病视觉概念,动态构建无冗余概念池
- 通过跨模态注意力融合概念语义,显著提升分类准确率
- 适合关注医疗图像持续学习与多模态融合的研究者
持续学习对适应动态临床环境的医学图像分类系统至关重要。多模态信息融合可显著提升图像类别持续学习性能。然而,现有方法虽利用文本模态信息,却仅依赖包含类别名的简单模板,忽视了更丰富的语义内容。为此,我们提出一种新框架,利用大语言模型(LLMs)生成的视觉概念作为判别性语义引导。方法通过基于相似性的过滤机制动态构建视觉概念池,避免冗余。为将概念融入持续学习过程,采用跨模态图像-概念注意力模块,并结合注意力损失。该模块可通过注意力机制利用相关视觉概念的语义知识,生成类别代表性融合特征以用于分类。在医学和自然图像数据集上的实验表明,本方法达到当前最优性能,验证了其有效性和优越性。代码将公开发布。
原文摘要 · Abstract (English)
Continual learning is essential for medical image classification systems to adapt to dynamically evolving clinical environments. The integration of multimodal information can significantly enhance continual learning of image classes. However, while existing approaches do utilize textual modality information, they solely rely on simplistic templates with a class name, thereby neglecting richer semantic information. To address these limitations, we propose a novel framework that harnesses visual concepts generated by large language models (LLMs) as discriminative semantic guidance. Our method dynamically constructs a visual concept pool with a similarity-based filtering mechanism to prevent redundancy. Then, to integrate the concepts into the continual learning process, we employ a cross-modal image-concept attention module, coupled with an attention loss. Through attention, the module can leverage the semantic knowledge from relevant visual concepts and produce class-representative fused features for classification. Experiments on medical and natural image datasets show our method achieves state-of-the-art performance, demonstrating the effectiveness and superiority of our method. We will release the code publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。