arXiv:2512.17396cs.CVcs.AI2025-12被引 5

构建大规模医学影像问答数据集,推动CT/MRI精准诊断研究

RadImageNet-VQA: A Large-Scale CT and MRI Dataset for Radiologic Visual Question Answering

  • 基于专家标注构建750万图像-问题对,覆盖97种病灶和8个解剖区域
  • 顶尖视觉语言模型在细粒度病理识别上仍表现不佳,尤其开放问答场景
  • 实验证明模型依赖图像而非文字线索,适合医疗AI可解释性研究

本文提出RadImageNet-VQA,一个面向CT和MRI影像的大型放射学视觉问答数据集。现有医学VQA数据集规模有限,多集中于X光或生物图示,且易受文本线索干扰。本数据集基于专家标注,包含750K张图像与750万组问题-答案样本,涵盖异常检测、解剖识别、病理识别三类任务,覆盖8个解剖区域及97种病理类型,支持开放式、封闭式和多选题。大量实验表明,当前先进视觉语言模型在细粒度病理识别上仍表现欠佳,尤其在开放式问答中,即使微调后也难突破。纯文本分析显示,无图像输入时模型性能接近随机,证实该数据集避免了语言捷径。完整数据集与基准已公开于https://huggingface.co/datasets/raidium/RadImageNet-VQA。

原文摘要 · Abstract (English)

In this work, we introduce RadImageNet-VQA, a large-scale dataset designed to advance radiologic visual question answering (VQA) on CT and MRI exams. Existing medical VQA datasets are limited in scale, dominated by X-ray imaging or biomedical illustrations, and often prone to text-based shortcuts. RadImageNet-VQA is built from expert-curated annotations and provides 750K images paired with 7.5M question-answer samples. It covers three key tasks - abnormality detection, anatomy recognition, and pathology identification - spanning eight anatomical regions and 97 pathology categories, and supports open-ended, closed-ended, and multiple-choice questions. Extensive experiments show that state-of-the-art vision-language models still struggle with fine-grained pathology identification, particularly in open-ended settings and even after fine-tuning. Text-only analysis further reveals that model performance collapses to near-random without image inputs, confirming that RadImageNet-VQA is free from linguistic shortcuts. The full dataset and benchmark are publicly available at https://huggingface.co/datasets/raidium/RadImageNet-VQA.

医学影像视觉问答数据集AI辅助诊断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。