arXiv:2501.16282eess.IVcs.AI2025-01被引 8

用轻量适配器提升多模态大模型对脑部疾病的诊断能力

Brain-Adapter: Enhancing Neurological Disorder Analysis with Adapter-Tuning Multimodal Large Language Models

  • 引入轻量瓶颈层,仅训练少量参数实现知识迁移
  • 结合CLIP策略对齐多模态数据,显著提升诊断准确率
  • 适合医学影像分析与临床辅助诊断研究者使用

理解脑部疾病对精准临床诊断与治疗至关重要。近年来,多模态大语言模型(MLLMs)为结合文本描述解析医学图像提供了新路径。然而,现有研究主要聚焦2D医学图像,忽视了3D图像中更丰富的空间信息,且单模态方法容易遗漏其他模态中的关键临床信息。为此,本文提出Brain-Adapter,通过引入额外的轻量瓶颈层,在不增加计算成本的前提下,学习并注入新知识至预训练模型。核心思想是利用轻量瓶颈层在少量参数下捕捉关键信息,并采用对比语言-图像预训练(CLIP)策略,将多模态数据对齐至统一表示空间。大量实验表明,该方法能有效融合多模态数据,显著提升诊断准确率,具备在真实诊断流程中应用的潜力。

原文摘要 · Abstract (English)

Understanding brain disorders is crucial for accurate clinical diagnosis and treatment. Recent advances in Multimodal Large Language Models (MLLMs) offer a promising approach to interpreting medical images with the support of text descriptions. However, previous research has primarily focused on 2D medical images, leaving richer spatial information of 3D images under-explored, and single-modality-based methods are limited by overlooking the critical clinical information contained in other modalities. To address this issue, this paper proposes Brain-Adapter, a novel approach that incorporates an extra bottleneck layer to learn new knowledge and instill it into the original pre-trained knowledge. The major idea is to incorporate a lightweight bottleneck layer to train fewer parameters while capturing essential information and utilize a Contrastive Language-Image Pre-training (CLIP) strategy to align multimodal data within a unified representation space. Extensive experiments demonstrated the effectiveness of our approach in integrating multimodal data to significantly improve the diagnosis accuracy without high computational costs, highlighting the potential to enhance real-world diagnostic workflows.

脑疾病分析多模态适配器大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。