arXiv:2506.04353cs.CVcs.AI2025-06被引 24

构建首个大规模胸片视觉问答基准,推动AI超越人类专家的影像诊断能力。

ReXVQA: A Large-scale Visual Question Answering Benchmark for Generalist Chest X-ray Understanding

  • 基于真实临床思维设计五类推理任务,覆盖病变存在、位置、否定判断等核心能力。
  • 最佳模型MedGemma准确率达83.24%,超过顶尖放射科医生77.27%的水平。
  • 首次通过真人读片对比证明AI在胸片理解上已超人类,适合研究通用医学AI者参考。

我们提出ReXVQA,目前规模最大、最全面的胸部放射学视觉问答(VQA)基准,包含约69.6万条问题与16万张胸片数据,涵盖训练、验证和测试集。不同于以往依赖模板生成的问题,ReXVQA引入多样且符合临床实际的任务集合,反映五大核心放射学推理技能:病变存在评估、位置分析、否定检测、鉴别诊断与几何推理。我们评估了八种先进的多模态大模型,包括MedGemma-4B-it、Qwen2.5-VL、Janus-Pro-7B和Eagle2-9B。表现最优的模型(MedGemma)整体准确率达83.24%。为缩小AI与临床专家之间的差距,我们开展包含3名放射科住院医师的综合人读研究,在200个随机样本上进行评估。结果显示,MedGemma达到83.84%准确率,优于最高水平的人类读者(77.27%),标志着AI在胸片解读中首次超越人类专家。人读研究揭示人工智能与人类专家之间存在显著性能差异:放射科医生间一致性高,而人类与AI之间则呈现更复杂的分歧模式。ReXVQA建立了评估通用放射学AI系统的全新标准,提供公开排行榜、细粒度划分、结构化解释及类别级分析。该基准为下一代能够模拟专家级临床推理的AI系统奠定基础。数据集将开源发布于https://huggingface.co/datasets/rajpurkarlab/ReXVQA。

原文摘要 · Abstract (English)

We present ReXVQA, the largest and most comprehensive benchmark for visual question answering (VQA) in chest radiology, comprising approximately 696,000 questions paired with 160,000 chest X-rays studies across training, validation, and test sets. Unlike prior efforts that rely heavily on template based queries, ReXVQA introduces a diverse and clinically authentic task suite reflecting five core radiological reasoning skills: presence assessment, location analysis, negation detection, differential diagnosis, and geometric reasoning. We evaluate eight state-of-the-art multimodal large language models, including MedGemma-4B-it, Qwen2.5-VL, Janus-Pro-7B, and Eagle2-9B. The best-performing model (MedGemma) achieves 83.24% overall accuracy. To bridge the gap between AI performance and clinical expertise, we conducted a comprehensive human reader study involving 3 radiology residents on 200 randomly sampled cases. Our evaluation demonstrates that MedGemma achieved superior performance (83.84% accuracy) compared to human readers (best radiology resident: 77.27%), representing a significant milestone where AI performance exceeds expert human evaluation on chest X-ray interpretation. The reader study reveals distinct performance patterns between AI models and human experts, with strong inter-reader agreement among radiologists while showing more variable agreement patterns between human readers and AI models. ReXVQA establishes a new standard for evaluating generalist radiological AI systems, offering public leaderboards, fine-grained evaluation splits, structured explanations, and category-level breakdowns. This benchmark lays the foundation for next-generation AI systems capable of mimicking expert-level clinical reasoning beyond narrow pathology classification. Our dataset will be open-sourced at https://huggingface.co/datasets/rajpurkarlab/ReXVQA

视觉问答医学影像大模型胸片分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。