医学视觉语言模型在真实临床推理中表现差,多模态理解能力严重不足。
The Illusion of Clinical Reasoning: A Benchmark Reveals the Pervasive Gap in Vision-Language Models for Clinical Competency
- 构建真实病例的骨骼关节基准测试,覆盖7类临床推理任务
- 顶尖模型在开放题上准确率仅60%以下,图像理解能力薄弱
- 医疗微调模型无优势,当前AI不适合独立临床决策
背景:基础模型快速融入临床与公共卫生,亟需超越狭义考试成绩的真正临床推理能力评估。现有基准多基于医考或精选病例,难以反映真实患者照护所需的多模态整合推理。方法:我们构建了骨骼与关节(B&J)基准,包含1,245个来自骨科与运动医学的真实病例问题,涵盖知识回忆、文本与图像解读、诊断生成、治疗规划及理由说明等7项临床推理任务。评估了11个视觉语言模型(VLMs)和6个大语言模型(LLMs),并与专家标注的基准答案对比。结果:不同任务间表现差距显著。顶尖模型在结构化选择题上准确率超90%,但在需多模态整合的开放题上准确率不足60%。VLMs在医学图像解读上存在明显局限,常出现严重文本驱动幻觉,忽略矛盾的视觉证据。值得注意的是,专门针对医疗领域微调的模型并未展现出持续优势。结论:当前人工智能模型尚未具备复杂多模态推理的临床胜任力,其安全部署应限于辅助性文本角色。未来在核心临床任务上的突破,依赖于多模态融合与视觉理解的根本性进展。
原文摘要 · Abstract (English)
Background: The rapid integration of foundation models into clinical practice and public health necessitates a rigorous evaluation of their true clinical reasoning capabilities beyond narrow examination success. Current benchmarks, typically based on medical licensing exams or curated vignettes, fail to capture the integrated, multimodal reasoning essential for real-world patient care. Methods: We developed the Bones and Joints (B&J) Benchmark, a comprehensive evaluation framework comprising 1,245 questions derived from real-world patient cases in orthopedics and sports medicine. This benchmark assesses models across 7 tasks that mirror the clinical reasoning pathway, including knowledge recall, text and image interpretation, diagnosis generation, treatment planning, and rationale provision. We evaluated eleven vision-language models (VLMs) and six large language models (LLMs), comparing their performance against expert-derived ground truth. Results: Our results demonstrate a pronounced performance gap between task types. While state-of-the-art models achieved high accuracy, exceeding 90%, on structured multiple-choice questions, their performance markedly declined on open-ended tasks requiring multimodal integration, with accuracy scarcely reaching 60%. VLMs demonstrated substantial limitations in interpreting medical images and frequently exhibited severe text-driven hallucinations, often ignoring contradictory visual evidence. Notably, models specifically fine-tuned for medical applications showed no consistent advantage over general-purpose counterparts. Conclusions: Current artificial intelligence models are not yet clinically competent for complex, multimodal reasoning. Their safe deployment should currently be limited to supportive, text-based roles. Future advancement in core clinical tasks awaits fundamental breakthroughs in multimodal integration and visual understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。