让医学视觉模型像医生一样看片、查文献,提升罕见病诊断准确率。
GAZE: Grounded Agentic Zero-shot Evaluation with Viewer-Level Tools and Literature Retrieval on Rare Brain MRI

- 模拟医生多次看片+查文献,调用工具与检索系统迭代分析
- 罕见病定位准确率从17%升至58%,诊断正确率达34.9%
- 适合医疗AI评估、罕见病诊断研究者参考
视觉语言模型通常单次推理完成图像理解,而放射科医生会反复观察影像并查阅文献。我们提出GAZE(基于锚定的智能体零样本评估)框架,使医学视觉语言模型可通过调用视图级工具(缩放、窗宽窗位、对比度调节、边缘检测)和两个由美国国家医学图书馆支持的检索工具(PubMed用于医学文献,Open-i用于影像),以迭代方式分析脑部MRI。框架输出结构化且符合模式校验,并记录完整工具调用轨迹以供审计。在涵盖281种罕见神经疾病、共906例脑MRI的NOVA基准上,GAZE在交并比0.3下实现58.2%的平均精度(mAP)用于病灶定位,在联合协议中诊断准确率为34.9%(评分包含描述、诊断、定位,无需任务微调)。在未使用任何工具前,结构化提示与模式验证已使基线(Gemini 2.0 Flash)的[email protected]从20.2提升至29.4。工具使用对罕见病改善更显著:三例以下疾病的定位成功率达58%(原17%),常见病(≥10例)达68%(原25%),工具调用次数与收益正相关(Gemini 3 Flash:Cohen's d=0.79,平均每例11.8次调用;Gemini 2.0 Flash:仅8.2%病例使用工具,无显著提升)。消融实验显示诊断与定位间存在模型依赖性权衡,强化了联合评估诊断、定位与描述的必要性。
原文摘要 · Abstract (English)
Vision-language models (VLMs) read an image and produce text in a single forward pass, whereas radiologists typically inspect an image several times and consult the literature before writing a report. We introduce GAZE (Grounded Agentic Zero-shot Evaluation), a framework that lets a medical VLM work in this iterative way by calling viewer-level tools (zoom, windowing, contrast, edge detection) and two retrieval tools backed by the U.S. National Library of Medicine (PubMed for medical literature, Open-i for radiological images), with structured outputs validated against a schema and full tool-call traces recorded for auditability. On NOVA, a benchmark of 906 brain MRI cases covering 281 rare neurological conditions, GAZE reaches 58.2 mean average precision (mAP) at intersection-over-union (IoU) 0.3 for lesion localisation and 34.9% Top-1 diagnostic accuracy under a joint protocol that scores captioning, diagnosis, and localisation from the image alone, without task-specific fine-tuning. Before any tool is used, structured prompting and schema-validated outputs already improve over the published Gemini 2.0 Flash baseline (20.2 to 29.4 [email protected]), so framework design is itself an experimental variable. Tool use helps rare pathologies disproportionately: the fraction of cases with IoU > 0.3 rises from 17% to 58% for diagnoses with three or fewer examples versus 25% to 68% for common conditions ($\geq$10 cases), with gains tracking engagement (Gemini 3 Flash: Cohen's d = 0.79, 11.8 tool calls per case; Gemini 2.0 Flash: tools used in 8.2% of cases, no significant benefit). Retrieval ablations additionally reveal a model-dependent trade-off in which gains in diagnosis can coincide with losses in localisation, reinforcing the case for joint evaluation of diagnosis, localisation, and captioning in medical VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。