arXiv:2509.07613cs.CV2025-09中稿 · MICAD 2025被引 1

用少量MRI数据高效微调视觉语言模型,提升阿尔茨海默病诊断能力。

Data-Efficient Fine-Tuning of Vision-Language Models for Diagnosis of Alzheimer's Disease

  • 将患者结构化数据转为合成报告,增强图文对齐。
  • 引入MMSE评分预测辅助任务,提升模型临床相关性。
  • 仅需1504张MRI即可超越训练量大20倍的模型,适合医疗小样本场景。

医学视觉语言模型(Med-VLMs)在报告生成和视觉问答等任务中表现优异,但仍存在不足:未能充分利用患者结构化数据,且缺乏临床诊断知识整合。现有模型多从头训练或在大规模二维图像-文本对上微调,计算开销大,且在三维医学影像上的效果受限于缺乏结构信息。为此,我们提出一种数据高效的微调流程,用于适配基于3D CT的Med-VLM以支持3D MRI,并应用于阿尔茨海默病(AD)诊断。系统引入两项关键创新:首先,将结构化元数据转换为合成报告,丰富文本输入以改善图像-文本对齐;其次,添加一个可学习的辅助标记,用于预测广泛使用的认知功能评估工具MMSE得分,该指标与AD严重程度密切相关,为微调提供额外监督。通过在图像和文本模态上应用轻量级提示微调,本方法在ADNI数据集上仅使用1,504张训练MRI即达到当前最佳性能,优于使用27,161张MRI训练的方法,并在OASIS-2和AIBL数据集上展现出强零样本泛化能力。代码已公开于https://github.com/CFQ666312/DEFT-VLM-AD。

原文摘要 · Abstract (English)

Medical vision-language models (Med-VLMs) have shown impressive results in tasks such as report generation and visual question answering, but they still face several limitations. Most notably, they underutilize patient metadata and lack integration of clinical diagnostic knowledge. Moreover, most existing models are typically trained from scratch or fine-tuned on large-scale 2D image-text pairs, requiring extensive computational resources, and their effectiveness on 3D medical imaging is often limited due to the absence of structural information. To address these gaps, we propose a data-efficient fine-tuning pipeline to adapt 3D CT-based Med-VLMs for 3D MRI and demonstrate its application in Alzheimer's disease (AD) diagnosis. Our system introduces two key innovations. First, we convert structured metadata into synthetic reports, enriching textual input for improved image-text alignment. Second, we add an auxiliary token trained to predict the mini-mental state examination (MMSE) score, a widely used clinical measure of cognitive function that correlates with AD severity. This provides additional supervision for fine-tuning. Applying lightweight prompt tuning to both image and text modalities, our approach achieves state-of-the-art performance on ADNI with only 1,504 training MRIs, outperforming methods trained on 27,161 MRIs, and shows strong zero-shot generalization on OASIS-2 and AIBL. Code is available at https://github.com/CFQ666312/DEFT-VLM-AD.

阿尔茨海默病小样本学习视觉语言模型医学AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。