通过知识蒸馏提升医学视觉语言模型的视觉对齐能力,减少幻觉输出。
Enhancing Medical Large Vision-Language Models via Alignment Distillation
- 用对比学习模型蒸馏视觉对齐知识,增强医学模型理解能力。
- 在报告生成与医学VQA任务中,性能与可解释性均显著提升。
- 适合需要高可信度医疗AI的临床场景应用。
医学大视觉语言模型(Med-LVLMs)在临床应用中表现良好,但常因视觉理解错位导致幻觉输出。本文识别出两个根本问题:视觉表征学习不足和视觉注意力对齐不佳。为此提出MEDALIGN,一种轻量级对齐蒸馏框架,将领域特定的对比语言-图像预训练(CLIP)模型中的视觉对齐知识迁移至Med-LVLMs。MEDALIGN引入两种蒸馏损失:基于视觉标记层级相似结构的空间感知对齐损失,以及引导注意力聚焦于诊断相关区域的注意力感知损失。在医学报告生成与医学视觉问答(VQA)基准上的大量实验表明,MEDALIGN持续提升了性能与可解释性,生成更贴近视觉输入的输出。
原文摘要 · Abstract (English)
Medical Large Vision-Language Models (Med-LVLMs) have shown promising results in clinical applications, but often suffer from hallucinated outputs due to misaligned visual understanding. In this work, we identify two fundamental limitations contributing to this issue: insufficient visual representation learning and poor visual attention alignment. To address these problems, we propose MEDALIGN, a simple, lightweight alignment distillation framework that transfers visual alignment knowledge from a domain-specific Contrastive Language-Image Pre-training (CLIP) model to Med-LVLMs. MEDALIGN introduces two distillation losses: a spatial-aware visual alignment loss based on visual token-level similarity structures, and an attention-aware distillation loss that guides attention toward diagnostically relevant regions. Extensive experiments on medical report generation and medical visual question answering (VQA) benchmarks show that MEDALIGN consistently improves both performance and interpretability, yielding more visually grounded outputs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。