CardioLens揭示了多模态大模型在真实心脏磁共振诊断中的巨大差距。
CardioLens: Revealing the Clinical Reality Gap of MLLMs via Multi-Sequence Cardiac MRI Evaluations

- 构建了基于医院私有数据的多序列心脏磁共振评估基准
- 发现现有模型在临床流程中性能显著下降,错误倾向常见异常类别
- 适合关注医疗AI临床落地与模型可靠性的研究者
多模态大语言模型在公开医学评测中表现优异,但现有评估仍难以反映真实临床场景,多依赖孤立输入和简化识别任务。本文提出CardioLens,一个基于私有医院档案、通过严谨报告到问答转换与验证流程构建的抗泄露评估基准,涵盖473,896张切片和13,494对已验证的问答对,覆盖4D cine、LGE、灌注及T2加权成像,评估心脏磁共振解读的三个阶段:图像理解、报告生成与疾病诊断。在24个先进MLLM上测试显示,模型整体表现不佳,且性能随真实临床工作流推进而持续下降。混淆分析揭示类别坍缩问题:模型倾向于默认为高发异常类别,而非区分临床差异性发现。对比随机、临床导向与数据驱动的切片选择策略(不同切片预算下),性能变化仅约1%,说明输入构造非主因;显式推理提示亦无法提升表现,反而使模型更保守。结果表明,当前MLLM距离可靠的心脏磁共振解读仍有巨大差距,临床决策需整合跨序列、视角与时间相位的分布证据。CardioLens为下一代面向真实临床部署的MLLM开发提供了坚实基准。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have shown strong performance on public medical benchmarks, yet existing evaluations often remain weak proxies for clinical use, relying on isolated inputs and simplified recognition-style tasks. We introduce CardioLens, a leakage-resistant evaluation testbed for multi-sequence Cardiovascular Magnetic Resonance (CMR), constructed from private hospital archives through a rigorous report-to-QA construction and verification pipeline. CardioLens contains 473,896 slices and 13,494 verified QA pairs across 4D Cine, LGE, perfusion, and T2-weighted imaging, and evaluates three stages of CMR interpretation: image understanding, report generation, and disease diagnosis. Across 24 state-of-the-art MLLMs, CardioLens reveals a substantial clinical reality gap: models perform poorly overall, with performance degrading along the real CMR workflow. Confusion analysis further shows a category-collapse failure mode, where models default to frequent abnormal categories rather than distinguishing clinically distinct findings. To rule out MLLM-compatible input construction as the primary cause, we compare random, clinically motivated, and data-driven slice selection protocols under different slice budgets; performance changes only marginally, typically by about 1%. Explicit reasoning prompts also fail to rescue performance, often making models more conservative rather than improving visual evidence use. These results show that current MLLMs remain far from reliable CMR interpretation, where clinical decisions require integrating distributed evidence across sequences, views, and temporal phases. CardioLens provides a clinically grounded testbed for developing next-generation MLLMs toward real-world clinical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。