用多图对齐提升医疗视觉语言模型的语义准确性,仅需10%数据达主流水平
ExGra-Med: Extended Context Graph Alignment for Medical Vision-Language Models
- 构建图像、回答和扩展描述的联合图对齐机制
- 仅用10%预训练数据即在VQA-RAD上提升20.13分
- 适合追求高效高质医疗多模态应用的研究者
当前领先的医疗多模态大模型(med-MLLMs)如LLaVA-Med和BioMedGPT主要依赖模型规模与数据量增长,训练以自回归目标为主。但我们发现该方法导致视觉-语言对齐薄弱,过度依赖昂贵的指令跟随数据。为此,我们提出ExGra-Med,一种新型多图对齐框架,联合对齐图像、指令响应与扩展描述的潜在空间,增强语义基础与跨模态一致性。为适配大模型(如LLaMA-7B),我们设计基于黑箱梯度估计的高效端到端训练方案,实现快速可扩展优化。实验表明,ExGra-Med仅用10%预训练数据即可达到LLaVA-Med性能,且在VQA-RAD上提升20.13%,在视觉聊天机器人和零样本分类任务中超越BioMedGPT与RadFM等强基线,展现其在医疗AI中高效高质量视觉语言融合的潜力。
原文摘要 · Abstract (English)
State-of-the-art medical multi-modal LLMs (med-MLLMs), such as LLaVA-Med and BioMedGPT, primarily depend on scaling model size and data volume, with training driven largely by autoregressive objectives. However, we reveal that this approach can lead to weak vision-language alignment, making these models overly dependent on costly instruction-following data. To address this, we introduce ExGra-Med, a novel multi-graph alignment framework that jointly aligns images, instruction responses, and extended captions in the latent space, advancing semantic grounding and cross-modal coherence. To scale to large LLMs (e.g., LLaMA-7B), we develop an efficient end-to-end training scheme using black-box gradient estimation, enabling fast and scalable optimization. Empirically, ExGra-Med matches LLaVA-Med's performance using just 10% of the pre-training data, achieving a 20.13% gain on VQA-RAD and approaching full-data performance. It also outperforms strong baselines like BioMedGPT and RadFM on visual chatbot and zero-shot classification tasks, demonstrating its promise for efficient, high-quality vision-language integration in medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。