首个评估医学多模态数据洞察发现能力的基准,揭示大模型短板并提出改进框架。
MedInsightBench: Evaluating Medical Analytics Agents Through Multi-Step Insight Discovery in Multimodal Medical Data
- 构建332个医学案例的多模态洞察评测基准,含精心设计的深度洞察标注。
- 现有大模型在多步推理与医学知识融合上表现不佳,平均得分低于40%。
- 提出MedInsightAgent框架,分三模块协同完成图像分析、洞察生成与追问设计。
在医学数据分析中,从复杂多模态数据中提取深层洞察对提升患者护理、提高诊断准确性和优化医疗运营至关重要。然而,目前缺乏专门用于评估大模型(LMMs)在多模态医学数据中发现医学洞察能力的高质量数据集。本文提出MedInsightBench,首个包含332个精心设计的医学案例的基准,每个案例均配有经过深思熟虑的洞察标注。该基准旨在评估大模型及智能体框架在分析多模态医学影像数据时的能力,包括提出相关问题、解释复杂发现以及综合生成可操作的洞察与建议。分析表明,现有大模型在MedInsightBench上的表现有限,主要归因于其在多步深度洞察提取和医学专业知识整合方面的挑战。因此,我们提出了MedInsightAgent——一种自动化医学数据分析智能体框架,由三个模块组成:视觉根因定位器、分析洞察智能体和后续问题生成器。在MedInsightBench上的实验揭示了普遍存在的挑战,并证明MedInsightAgent能显著提升通用大模型在医学洞察发现中的性能。
原文摘要 · Abstract (English)
In medical data analysis, extracting deep insights from complex, multi-modal datasets is essential for improving patient care, increasing diagnostic accuracy, and optimizing healthcare operations. However, there is currently a lack of high-quality datasets specifically designed to evaluate the ability of large multi-modal models (LMMs) to discover medical insights. In this paper, we introduce MedInsightBench, the first benchmark that comprises 332 carefully curated medical cases, each annotated with thoughtfully designed insights. This benchmark is intended to evaluate the ability of LMMs and agent frameworks to analyze multi-modal medical image data, including posing relevant questions, interpreting complex findings, and synthesizing actionable insights and recommendations. Our analysis indicates that existing LMMs exhibit limited performance on MedInsightBench, which is primarily attributed to their challenges in extracting multi-step, deep insights and the absence of medical expertise. Therefore, we propose MedInsightAgent, an automated agent framework for medical data analysis, composed of three modules: Visual Root Finder, Analytical Insight Agent, and Follow-up Question Composer. Experiments on MedInsightBench highlight pervasive challenges and demonstrate that MedInsightAgent can improve the performance of general LMMs in medical data insight discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。