用指南图页增强眼科问诊,让AI回答更精准可信。
Guideline-grounded retrieval-augmented generation for ophthalmic clinical decision support
- 将指南页当证据单元,直接检索图像保全图表布局
- 硬数据集上准确率提升至65.76%,比GPT-5.2高10.4%
- 适合需要严格依据指南的临床AI系统研发者
本文提出Oph-Guid-RAG,一种面向眼科临床问答与决策支持的多模态视觉RAG系统。将每页指南视为独立证据单元,直接检索页面图像,保留表格、流程图和版式信息。设计可控检索框架,通过路由与过滤机制选择性引入外部证据,减少噪声。系统整合查询分解、重写、检索、重排序与多模态推理,输出可追溯并附指南页引用。在HealthBench上采用医生评分协议评估,硬子集上整体得分从0.2969提升至0.3861(+0.0892,+30.0%),准确率由0.5956升至0.6576(+0.0620,+10.4%)。相比GPT-5.4,准确率提升达+0.1289(+24.4%),表明该方法在需精确证据推理的难题上更具优势。消融实验显示,重排序、路由与检索设计对稳定性能至关重要,尤其在困难场景下。结果表明,结合视觉检索与可控推理可提升临床AI的证据锚定性与鲁棒性,但仍需进一步完善。
原文摘要 · Abstract (English)
In this work, we propose Oph-Guid-RAG, a multimodal visual RAG system for ophthalmology clinical question answering and decision support. We treat each guideline page as an independent evidence unit and directly retrieve page images, preserving tables, flowcharts, and layout information. We further design a controllable retrieval framework with routing and filtering, which selectively introduces external evidence and reduces noise. The system integrates query decomposition, query rewriting, retrieval, reranking, and multimodal reasoning, and provides traceable outputs with guideline page references. We evaluate our method on HealthBench using a doctor-based scoring protocol. On the hard subset, our approach improves the overall score from 0.2969 to 0.3861 (+0.0892, +30.0%) compared to GPT-5.2, and achieves higher accuracy, improving from 0.5956 to 0.6576 (+0.0620, +10.4%). Compared to GPT-5.4, our method achieves a larger accuracy gain of +0.1289 (+24.4%). These results show that our method is more effective on challenging cases that require precise, evidence-based reasoning. Ablation studies further show that reranking, routing, and retrieval design are critical for stable performance, especially under difficult settings. Overall, we show how combining visionbased retrieval with controllable reasoning can improve evidence grounding and robustness in clinical AI applications,while pointing out that further work is needed to be more complete.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。