构建多轮多模态医疗推理评测框架,推动AI更贴近真实临床决策。
MedAtlas: Evaluating LLMs for Multi-Round, Multi-Task Medical Reasoning Across Diverse Imaging Modalities and Clinical Text
- 设计多轮对话+多模态图像交互的临床推理任务
- 在真实诊断流程上,现有模型表现显著落后于专家水平
- 适合医学AI研发者与临床推理系统评估者使用
人工智能在临床决策中展现出巨大潜力,但开发能适应多样化现实场景并完成复杂诊断推理的模型仍面临重大挑战。现有医学多模态基准通常仅限于单图像、单轮任务,缺乏多模态医学影像整合,也未能体现临床实践中固有的纵向性与多模态交互特性。为填补这一空白,我们提出MedAtlas——一个新型基准框架,用于评估大语言模型在真实医疗推理任务中的表现。MedAtlas具备四大特征:多轮对话、多模态医学影像交互、多任务集成和高临床保真度。支持四项核心任务:开放式多轮问答、封闭式多轮问答、多图像联合推理和综合疾病诊断。每例均源自真实诊断流程,融合文本病史与多种影像模态(包括CT、MRI、PET、超声、X-ray)的时间交互,要求模型在图像与临床文本间进行深度整合推理。所有任务均提供专家标注的黄金标准。此外,我们提出两个新评估指标:回合链准确率与错误传播抵抗能力。现有多模态模型的基准测试结果揭示了在多阶段临床推理上的显著性能差距。MedAtlas为推进稳健可信的医疗AI发展提供了具有挑战性的评估平台。
原文摘要 · Abstract (English)
Artificial intelligence has demonstrated significant potential in clinical decision-making; however, developing models capable of adapting to diverse real-world scenarios and performing complex diagnostic reasoning remains a major challenge. Existing medical multi-modal benchmarks are typically limited to single-image, single-turn tasks, lacking multi-modal medical image integration and failing to capture the longitudinal and multi-modal interactive nature inherent to clinical practice. To address this gap, we introduce MedAtlas, a novel benchmark framework designed to evaluate large language models on realistic medical reasoning tasks. MedAtlas is characterized by four key features: multi-turn dialogue, multi-modal medical image interaction, multi-task integration, and high clinical fidelity. It supports four core tasks: open-ended multi-turn question answering, closed-ended multi-turn question answering, multi-image joint reasoning, and comprehensive disease diagnosis. Each case is derived from real diagnostic workflows and incorporates temporal interactions between textual medical histories and multiple imaging modalities, including CT, MRI, PET, ultrasound, and X-ray, requiring models to perform deep integrative reasoning across images and clinical texts. MedAtlas provides expert-annotated gold standards for all tasks. Furthermore, we propose two novel evaluation metrics: Round Chain Accuracy and Error Propagation Resistance. Benchmark results with existing multi-modal models reveal substantial performance gaps in multi-stage clinical reasoning. MedAtlas establishes a challenging evaluation platform to advance the development of robust and trustworthy medical AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。