评测20个视觉大模型在神经影像诊断中的表现,发现肿瘤分类最准,多发性硬化仍难搞定。
NeuroVLM-Bench: Evaluation of Vision-Enabled Large Language Models for Clinical Reasoning in Neurological Disorders
- 构建多模态医学影像评测基准,要求模型同时输出诊断、分型、成像方式等结构化结果。
- 肿瘤分类准确率最高,多发性硬化和罕见异常仍难识别,少样本提示可提升性能但增加成本。
- 开源模型MedGemma-1.5-4B表现亮眼,接近闭源模型零样本效果且输出格式完美。
近期多模态大语言模型为基于图像的临床决策支持带来新可能,但其在神经影像中的可靠性与实际权衡尚不明确。本文针对多发性硬化、卒中、脑肿瘤、其他异常及正常对照的MRI与CT数据集,开展全面基准评测,要求模型同步生成诊断、亚型、成像模态、特殊序列和解剖平面。评估涵盖判别分类(含拒答)、校准性、结构化输出有效性及计算效率四个维度。采用多阶段框架控制选择偏差。对20个前沿多模态模型的评测显示,成像属性(如模态、平面)已基本解决,而诊断推理尤其是亚型预测仍具挑战;肿瘤分类最可靠,卒中可部分解决,多发性硬化和罕见异常仍困难。少样本提示能提升部分模型性能,但增加令牌消耗、延迟与成本。Gemini-2.5-Pro与GPT-5-Chat整体诊断性能最强,Gemini-2.5-Flash在效率-性能间平衡最优。开源模型中,MedGemma-1.5-4B表现最佳:少样本提示下接近多个闭源模型的零样本性能,且保持完全结构化输出。研究结果为多模态大模型在神经影像中的应用提供实用参考。
原文摘要 · Abstract (English)
Recent advances in multimodal large language models enable new possibilities for image-based decision support. However, their reliability and operational trade-offs in neuroimaging remain insufficiently understood. We present a comprehensive benchmarking study of vision-enabled large language models for 2D neuroimaging using curated MRI and CT datasets covering multiple sclerosis, stroke, brain tumors, other abnormalities, and normal controls. Models are required to generate multiple outputs simultaneously, including diagnosis, diagnosis subtype, imaging modality, specialized sequence, and anatomical plane. Performance is evaluated across four directions: discriminative classification with abstention, calibration, structured-output validity, and computational efficiency. A multi-phase framework ensures fair comparison while controlling for selection bias. Across twenty frontier multimodal models, the results show that technical imaging attributes such as modality and plane are nearly solved, whereas diagnostic reasoning, especially subtype prediction, remains challenging. Tumor classification emerges as the most reliable task, stroke is moderately solvable, while multiple sclerosis and rare abnormalities remain difficult. Few-shot prompting improves performance for several models but increases token usage, latency, and cost. Gemini-2.5-Pro and GPT-5-Chat achieve the strongest overall diagnostic performance, while Gemini-2.5-Flash offers the best efficiency-performance trade-off. Among open-weight architectures, MedGemma-1.5-4B demonstrates the most promising results, as under few-shot prompting, it approaches the zero-shot performance of several proprietary models, while maintaining perfect structured output. These findings provide practical insights into performance, reliability, and efficiency trade-offs, supporting standardized evaluation of multimodal LLMs in neuroimaging.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。