arXiv:2606.28556cs.AI2026-06中稿 · KDD

首个面向医学多模态对话的基准,评测模型在真实影像下的诊疗安全性。

IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations

论文配图:IMCBench: A benchmark for multimodal LLMs in Image-grounded Medical Conversations
图 1 · 摘自论文原文
  • 构建含真实医学影像与病历的多轮对话数据集,模拟医患交互场景。
  • 模型最高得分3.61,但恶性与罕见病安全性能下降0.27,暴露风险。
  • 视觉信息和病历上下文对安全决策至关重要,强模型更善用图像特征。

近年来,大语言模型与视觉-语言模型的发展推动了多模态推理在临床决策支持与分诊中的应用。然而现有医疗AI基准存在碎片化问题:部分支持多轮对话但缺乏图像,另一些虽提供多模态输入却仅限单轮问答。为弥补此缺口,我们提出IMCBench——一个基于真实临床图像、结合合成患者档案的多轮医学对话基准,用于模拟真实医患互动。每轮对话从安全性、准确性及诊断不确定性使用三个维度评估。我们对八种前沿多模态模型(来自Claude、GPT、Nova、Llama四类)进行评测,采用经专家标注校准的LLM-as-Jury评分体系,满分5分。结果显示,Claude Opus 4.6得分最高(3.61),次之为Claude Sonnet 4.6(3.30)和GPT-5.2(3.29),但无模型在所有维度领先;在恶性与罕见病情境下,安全性均下降0.27。消融实验表明,移除视觉输入或电子病历(EHR)上下文时,安全性平均分别下降0.18与0.23,且更强模型更有效利用视觉特征。结果表明,准确描述不等于安全引导,亟需多维评价框架。

原文摘要 · Abstract (English)

Recent advances in large language models and vision-language models have enabled reasoning over multimodal data, offering opportunities for clinical applications such as decision support and triaging. However, existing medical AI benchmarks are fragmented: some support multi-turn dialogues but lack images, while others provide multimodal inputs but focus on single-turn QA tasks. To address this gap, we introduce IMCBench, an image-grounded, multi-turn medical conversation benchmark that pairs real, publicly available clinical images with synthetic patient profiles to simulate realistic patient-clinician interactions. Each conversation is evaluated across three clinical dimensions: safety, accuracy, and appropriate use of uncertainty in diagnosis. We benchmark eight multimodal frontier models across four model families (Claude, GPT, Nova, and Llama), scoring each on a 1-5 scale using LLM-as-Jury scoring calibrated against expert clinician annotations. Our results show that Claude Opus 4.6 achieves the highest overall score (3.61), followed by Claude Sonnet 4.6 (3.30) and GPT-5.2 (3.29), though no model dominates all dimensions and safety degrades for both malignant and rare conditions ($Δ$ = -0.27 each). Ablation studies further reveal that both visual input and EHR context contribute to safe guidance (safety drops of 0.18 and 0.23 on average when each is removed), with stronger models leveraging visual features more effectively. Together, these findings demonstrate that accurate clinical description does not guarantee safe patient guidance, motivating the need for multi-dimensional evaluation frameworks in medical AI.

多模态医疗AI对话系统基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。