arXiv:2506.00855cs.AI2025-06被引 2

基于开源医学教材构建多模态医疗问答基准,评估AI在临床任务中的表现。

MedBookVQA: A Systematic and Comprehensive Medical Benchmark Derived from Open-Access Book

  • 从开放获取医学教材中自动提取图文并构建对齐数据集
  • 生成5000个涵盖5类临床任务的医学问题,覆盖42种影像模态
  • 首次系统评估多类模型在解剖结构与专科任务上的性能差异

通用医疗人工智能(GMAI)借助多模态大语言模型(MLLMs)快速发展,有望缓解医疗人力短缺与成本上升等长期挑战。与此同时,系统性评估基准的建设成为必要前提。尽管医学教材是宝贵的知识来源,其在基准构建中的潜力尚未被充分挖掘。本文提出MedBookVQA,一个源自开放获取医学教材的系统化、综合性多模态基准。我们设计标准化流程,实现医学图像的自动化提取并与对应文本上下文精准对齐。基于该数据集,生成5000个临床相关问题,涵盖模态识别、疾病分类、解剖定位、症状诊断及手术操作等任务。采用多层次标注体系,按医学影像模态(42类)、人体解剖结构(125个)、临床专科(31个部门)进行分层分类,支持跨亚领域精细分析。我们评估了包括专有、开源、医疗专用及推理型在内的多种MLLM,揭示不同任务类型与模型类别间显著性能差异。研究发现当前GMAI系统存在关键能力缺口,确立教材衍生多模态基准作为关键技术评估工具的重要性。MedBookVQA确立了以教材为基础的基准范式,暴露GMAI系统的局限性,并提供按解剖结构与专科划分的性能指标。

原文摘要 · Abstract (English)

The accelerating development of general medical artificial intelligence (GMAI), powered by multimodal large language models (MLLMs), offers transformative potential for addressing persistent healthcare challenges, including workforce deficits and escalating costs. The parallel development of systematic evaluation benchmarks emerges as a critical imperative to enable performance assessment and provide technological guidance. Meanwhile, as an invaluable knowledge source, the potential of medical textbooks for benchmark development remains underexploited. Here, we present MedBookVQA, a systematic and comprehensive multimodal benchmark derived from open-access medical textbooks. To curate this benchmark, we propose a standardized pipeline for automated extraction of medical figures while contextually aligning them with corresponding medical narratives. Based on this curated data, we generate 5,000 clinically relevant questions spanning modality recognition, disease classification, anatomical identification, symptom diagnosis, and surgical procedures. A multi-tier annotation system categorizes queries through hierarchical taxonomies encompassing medical imaging modalities (42 categories), body anatomies (125 structures), and clinical specialties (31 departments), enabling nuanced analysis across medical subdomains. We evaluate a wide array of MLLMs, including proprietary, open-sourced, medical, and reasoning models, revealing significant performance disparities across task types and model categories. Our findings highlight critical capability gaps in current GMAI systems while establishing textbook-derived multimodal benchmarks as essential evaluation tools. MedBookVQA establishes textbook-derived benchmarking as a critical paradigm for advancing clinical AI, exposing limitations in GMAI systems while providing anatomically structured performance metrics across specialties.

医疗AI多模态评测基准医学图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。