构建可解释的胸部X光问答数据集,支持病灶定位与临床推理。
VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning
- 基于放射科医生标注的边界框和推理说明构建多任务问答数据集
- 在4,394张图像上覆盖17,597个问题,正负样本比例平衡为41.7%:58.3%
- 支持病变定位,使模型回答更可信,适合医疗AI可解释性研究
我们提出VinDr-CXR-VQA,一个大规模胸部X光可解释医学视觉问答(Med-VQA)数据集,具备空间定位能力。该数据集包含4,394张图像上的17,597个问答对,每对均配有放射科医生验证的边界框和临床推理说明。问题类型涵盖六类诊断意图:位置、内容、是否存在、数量、哪个、是/否,覆盖多样临床需求。为提升可靠性,构建了41.7%阳性与58.3%阴性样本的平衡分布,有效缓解正常病例中的幻觉问题。以MedGemma-4B-it为基准进行评测,性能显著提升(F1=0.624,较基线+11.8%),同时实现病灶定位。该数据集及评估工具已公开发布于huggingface.co/datasets/Dangindev/VinDR-CXR-VQA,旨在推动可复现且临床可落地的Med-VQA研究。
原文摘要 · Abstract (English)
We present VinDr-CXR-VQA, a large-scale chest X-ray dataset for explainable Medical Visual Question Answering (Med-VQA) with spatial grounding. The dataset contains 17,597 question-answer pairs across 4,394 images, each annotated with radiologist-verified bounding boxes and clinical reasoning explanations. Our question taxonomy spans six diagnostic types-Where, What, Is there, How many, Which, and Yes/No-capturing diverse clinical intents. To improve reliability, we construct a balanced distribution of 41.7% positive and 58.3% negative samples, mitigating hallucinations in normal cases. Benchmarking with MedGemma-4B-it demonstrates improved performance (F1 = 0.624, +11.8% over baseline) while enabling lesion localization. VinDr-CXR-VQA aims to advance reproducible and clinically grounded Med-VQA research. The dataset and evaluation tools are publicly available at huggingface.co/datasets/Dangindev/VinDR-CXR-VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。