构建首个3D脑MRI视觉问答大基准,支持多病种临床推理评估
NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding

- 基于12个数据集的56,953对问答,覆盖全脑3D影像与多病种临床场景
- 采用图像依赖性验证机制,将文本捷径准确率降至44.6%
- 专为医生与模型设计,适合评估临床智能辅助系统能力
我们提出NeuroQA,一个大规模的3D脑磁共振成像(MRI)视觉问答基准,包含来自12,977名受试者、跨越12个数据集的56,953个问答对。研究对象年龄跨度5-104岁,覆盖阿尔茨海默病、帕金森病、肿瘤、白质病变和神经发育五大临床领域。不同于以往基于2D切片或依赖窄诊断标签的医学视觉问答工作,NeuroQA为每个条目配以完整的3D体积数据,评估11种临床相关推理能力,涵盖是/否、多选和开放式题型。在203个模板中,131个为图像可解(可通过三平面视图回答),72个为图像信息型(真实答案来自定量容积测量或临床工具)。为消除纯文本捷径,采用答案分布精炼策略,使闭式题型的纯文本准确率从>80%降至44.6%;图像必要性通过配套发布图像定位协议独立评估。使用38条规则的确定性流水线及两轮专家评审,所有问答对均经自由表面分析(FreeSurfer)测量、元数据或放射科报告字段验证,各模板间无同一受试者矛盾。我们进行了临床医生评估:两位医生独立在三平面查看器上评估100个冻结测试项。在闭式题型(是/否+多选)测试公开项上,最佳零样本视觉语言模型和监督3D CNN基线分别达到47.5%和43.7%准确率,均低于49.4%的文本主导模板基准。NeuroQA采用双层发布机制:开放数据集提供公共问答对与可复现生成脚本,受限数据集通过数据使用协议(DUAs)管理;同时提供受试者级划分、保留私有测试集与在线排行榜。
原文摘要 · Abstract (English)
We present NeuroQA, a large-scale benchmark for visual question answering in 3D brain magnetic resonance imaging (MRI), with 56,953 QA pairs from 12,977 subjects across 12 datasets. It spans ages 5-104 and five clinical domains: Alzheimer's, Parkinson's, tumors, white matter disease, and neurodevelopment. Unlike prior medical Visual Question Answering (VQA) efforts that operate on 2D slices or rely on narrow diagnostic labels, NeuroQA pairs every item with a full 3D volume. It evaluates 11 clinically grounded reasoning skills across Yes/No, multiple-choice, and open-ended formats. Of the 203 templates, 131 are image-grounded (answerable from a 3-plane viewer) and 72 are image-informed (ground truth from quantitative volumetry or clinical instruments). To remove text-only shortcuts, we apply answer-distribution refinement, reducing closed-format text-only accuracy from $>$80% to 44.6%; image necessity is assessed separately through an image-grounding protocol released with the benchmark. A 38-rule deterministic pipeline and two rounds of expert review verify every QA pair against FreeSurfer measurements, metadata, or radiology report fields, with zero same-subject contradictions across templates. We conduct a clinician evaluation in which two clinicians independently assess 100 frozen test items on a three-plane viewer. On closed-format (Yes/No + multiple-choice) test-public items, the best zero-shot vision-language model and a supervised 3D CNN baseline reach 47.5% and 43.7% accuracy respectively, both below the 49.4% text-only majority-template floor. NeuroQA adopts a two-tier release with public QA pairs for open-access datasets and reproducible generation scripts for datasets restricted by data use agreements (DUAs), plus subject-level splits, a held-out private test set, and an online leaderboard.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。