构建分阶段3D CT肿瘤诊断评估基准,揭示AI模型真实短板。
DeepTumorVQA: A Hierarchical 3D CT Benchmark for Stage-Wise Evaluation of Medical VLMs and Tool-Augmented Agents

- 按诊断流程分四阶段:识别、测量、视觉推理、医学推理,逐级评估
- 30种模型测试发现测量不准是主要瓶颈,工具辅助可显著提升表现
- 支持工具调用环境,适合研究医疗AI代理与多步推理的开发者
医学视觉语言模型(VLMs)和AI代理在临床图像分析中进展显著,但现有医学视觉问答(VQA)基准将模型能力压缩为单一准确率,掩盖了失败位置与原因。本文提出DeepTumorVQA,一个分层基准,遵循肿瘤诊断中的多阶段证据链,将3D CT推理分解为四个阶段:识别、测量、视觉推理和医学推理。高层问题可独立评分,其真实答案链基于底层基础项定义。该基准包含476,000个问题,覆盖9,262个3D CT影像中的42种临床亚型。除直接推理模式外,还提供工具交互环境,允许模型调用分割模型、测量程序和医学知识模块。对30种模型配置的评估显示,可靠定量测量是主要瓶颈,使后续视觉与医学推理更困难;而工具增强可显著缓解此问题。当工具可用时,利用医学知识与工具进行图像推理成为新挑战。我们进一步证明,DeepTumorVQA提供的真实分步工具使用轨迹可有效监督代理,减少工具使用与推理错误。从识别到测量再到视觉与医学推理的阶段演进,为未来医学VLM与AI代理研究提供了明确路线图。所有数据与代码已开源于https://github.com/Schuture/DeepTumorVQA。
原文摘要 · Abstract (English)
Medical vision-language models (VLMs) and AI agents have made significant progress in learning to analyze and reason about clinical images. However, existing medical visual question answering (VQA) benchmarks collapse model capabilities into a single accuracy score, obscuring where and why models fail. We propose DeepTumorVQA, a hierarchical benchmark that follows the multi-stage evidence chain in tumor diagnosis and decomposes 3D CT reasoning into four stages: recognition, measurement, visual reasoning, and medical reasoning. Higher-level questions remain independently scorable, while their ground-truth evidence chains are defined over lower-level primitives. The benchmark contains 476K questions across 42 clinical subtypes on 9,262 3D CT volumes. In addition to a direct reasoning mode for VLMs, DeepTumorVQA provides tool-interaction environments for agent evaluation, where a model can call external tools, including segmentation models, measurement programs, and medical knowledge modules, before answering the question. Evaluating over 30 model configurations, we find that reliable quantitative measurement is the primary bottleneck, making later-stage visual and medical reasoning harder for VLMs, while tool augmentation substantially mitigates this issue. When tools are available, leveraging medical knowledge and tools to reason about medical images becomes a new challenge. We further show that ground-truth step-by-step tool-use traces from DeepTumorVQA can supervise agents and reduce tool-use and reasoning failures. This stage-wise progression from recognition to measurement to visual and medical reasoning provides a concrete roadmap for future medical VLM and AI agent studies. All data and code are released at https://github.com/Schuture/DeepTumorVQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。