构建可解释的眼科视觉问答数据集,定位病变位置提升诊断可信度。
Towards Clinically Interpretable Ophthalmic VQA via Spatially-Grounded Lesion Evidence

- 用ETDRS网格精确定位15,595个病变,实现空间对齐的可视化证据
- 生成72,706个问题,涵盖四种题型,覆盖临床常见问答场景
- 引入双指标评估,证明显式空间证据能提升模型准确率与可解释性
视觉问答(VQA)在眼科临床支持中潜力巨大,因眼底照片是诊断关键。但现有眼科VQA基准主要关注答案准确率,忽视了临床所需的显式视觉证据。本文提出FundusGround,一个面向可解释眼科VQA的新基准,包含10,719张眼底图像和15,595个逐图像精细标注的病变。所有病变均通过早期糖尿病视网膜病变研究(ETDRS)网格进行空间定位,标准化映射至九个临床相关视网膜区域。基于此结构化病变证据,生成72,706个问题,涵盖开放、封闭、单选和多选四种格式。我们采用双指标(答案准确率与病灶级推理)评测多个通用及医疗视觉-语言大模型。实验表明,引入病灶级视觉证据能持续提升模型性能与透明度,凸显显式空间定位对可靠且可解释眼科VQA的必要性。
原文摘要 · Abstract (English)
Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential for diagnosis. However, ophthalmic VQA benchmarks primarily emphasize answer accuracy, neglecting the explicit visual evidence necessary for clinical interpretability. In this work, we introduce FundusGround, a new benchmark for clinically interpretable ophthalmic VQA with spatially-grounded lesion evidence. Specifically, we propose a three-stage pipeline that collects 10,719 fundus images with 15,595 image-level meticulously annotated lesions. To ensure anatomical consistency and clinical validity, all lesions are spatially localized using the Early Treatment Diabetic Retinopathy Study (ETDRS) grid, enabling standardized mapping to nine clinically meaningful retinal regions. Built upon this structured lesion evidence, 72,706 questions are then generated spanning four formats: open-ended, closed-ended, single-choice, and multiple-choice. We further benchmark multiple general- and medical- large vision-language models using dual metrics for answer accuracy and lesion-level reasoning. The experiments demonstrate that incorporating lesion-level visual evidence consistently improves model performance and transparency, highlighting the necessity of explicit spatial grounding for reliable and explainable ophthalmic VQA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。