arXiv:2606.27500cs.CVcs.CL2026-06

构建医疗视觉语言模型新基准,提升医学多模态模型的可靠性与可复现性。

Aloe-Vision: Robust Vision-Language Models for Healthcare

论文配图:Aloe-Vision: Robust Vision-Language Models for Healthcare
图 1 · 摘自论文原文
  • 整合医疗与通用数据构建高质量多模态训练集,支持模型微调。
  • 推出7B与72B双规模模型,性能媲美顶尖水平且保持通用能力。
  • 设计防污染视觉评测集,适用于临床场景的可靠评估。

面向医疗领域的大型视觉语言模型(LVLM)虽具临床应用潜力,但受限于高质量多模态数据稀缺、安全性关键场景下鲁棒性不足,以及评估基准狭窄且易受污染。为此,本文提出Aloe-Vision-Data,一个大规模、经质量筛选的混合数据集,融合医疗与通用领域,涵盖多模态与纯文本来源,可直接用于模型微调。基于此数据集,我们训练并开源了Aloe-Vision系列医疗LVLM,包含7B和72B两个规模版本,提供完整权重、训练配方与数据。综合评测表明,高质量训练混合能生成平衡性能的模型,在不牺牲通用能力前提下显著优于基线模型,达到与当前先进方法相当的水平。为支持可靠评估,我们构建CareQA-Vision,源自西班牙医/护专业资格考试(MIR与EIR),包含全新视觉题项,有效降低污染风险。最后实验揭示现有LVLM仍对对抗性与误导性输入敏感,凸显临床部署中的可靠性挑战。

原文摘要 · Abstract (English)

Large Vision-Language Models (LVLMs) specialized in healthcare are emerging as a promising research direction due to their potential impact in clinical and biomedical applications. However, progress is constrained by the scarcity of high-quality medical multimodal data, concerns about robustness in safety-critical settings, and the narrow and potentially contaminated evaluation benchmarks that limit reliable assessment. To address these issues, the field requires state-of-the-art solutions to be fully open and reproducible systems in which all components can be inspected, evaluated, and improved. This work introduces Aloe-Vision-Data, a large-scale, quality-filtered mixture which integrates both medical and general domains across multimodal and text-only sources, designed for direct use in model fine-tuning. Building on this dataset, we train the Aloe-Vision family of medical LVLMs, openly released with full weights, training recipes and data, in two scales (7B and 72B). Through comprehensive benchmarking, we demonstrate that high quality training mixtures produce balanced LVLMs which yield significant gains over the baseline models without compromising general capabilities, achieving competitive performance with respect to state-of-the-art alternatives. To support reliable evaluation, we introduce CareQA-Vision, a carefully curated vision benchmark derived from MIR and EIR exams, the residency entrance exams for medical and nursing specialists in Spain, offering novel vision questions with low likelihood of contamination. Finally, we show that current LVLMs remain vulnerable to adversarial and misleading inputs, underscoring reliability challenges in clinical contexts.

医疗AI视觉语言模型数据集可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。