arXiv:2506.11147cs.CV2025-06NeurIPS被引 24

构建首个支持多时序分析的3D放射科VQA数据集,推动医学影像智能诊断

3D-RAD: A Comprehensive 3D Radiology Med-VQA Dataset with Multi-Temporal Analysis and Diverse Diagnostic Tasks

  • 设计六类诊断任务,覆盖异常检测到纵向时间分析
  • 包含13.6万条专家标注样本,支持开闭式问答和复杂推理
  • 揭示现有模型在3D多时序任务中表现差,适合研究医疗多模态AI者

医学视觉问答(Med-VQA)在临床决策支持中潜力巨大,但现有研究多集中于2D影像且任务类型单一。本文提出3D-RAD,一个基于放射科CT扫描的大规模3D Med-VQA数据集。该数据集涵盖六类多样化任务:异常检测、图像观察、医学计算、存在性检测、静态时序诊断和纵向时序诊断。支持开闭式问题,引入计算任务与多阶段时序分析等复杂推理挑战,实现全面基准测试。大量评估表明,现有视觉语言模型(尤其是医学专用模型)在多时序任务上泛化能力有限,凸显真实3D诊断推理的挑战。为推动后续发展,我们发布高质量训练集3D-RAD-T,含136,195个专家对齐样本,实验证明在该数据集上微调可显著提升模型性能。数据集与代码已公开,旨在促进多模态医疗AI研究,建立3D医学视觉理解的坚实基础。

原文摘要 · Abstract (English)

Medical Visual Question Answering (Med-VQA) holds significant potential for clinical decision support, yet existing efforts primarily focus on 2D imaging with limited task diversity. This paper presents 3D-RAD, a large-scale dataset designed to advance 3D Med-VQA using radiology CT scans. The 3D-RAD dataset encompasses six diverse VQA tasks: anomaly detection, image observation, medical computation, existence detection, static temporal diagnosis, and longitudinal temporal diagnosis. It supports both open- and closed-ended questions while introducing complex reasoning challenges, including computational tasks and multi-stage temporal analysis, to enable comprehensive benchmarking. Extensive evaluations demonstrate that existing vision-language models (VLMs), especially medical VLMs exhibit limited generalization, particularly in multi-temporal tasks, underscoring the challenges of real-world 3D diagnostic reasoning. To drive future advancements, we release a high-quality training set 3D-RAD-T of 136,195 expert-aligned samples, showing that fine-tuning on this dataset could significantly enhance model performance. Our dataset and code, aiming to catalyze multimodal medical AI research and establish a robust foundation for 3D medical visual understanding, are publicly available at https://github.com/Tang-xiaoxiao/3D-RAD.

医学视觉问答3D影像多时序分析数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。