arXiv:2604.08322cs.CV2026-04被引 3

用公开数据训练眼科影像理解大模型,提升诊断推理能力。

Fundus-R1: Training a Fundus-Reading MLLM with Knowledge-Aware Reasoning on Public Data

  • 用检索增强生成法自动生成带眼科知识的推理链条。
  • 在三个基准上超越通用模型和未用推理链的版本。
  • 适合想用公共数据做医疗视觉语言研究的人。

眼底成像(如CFP、OCT和UWF)对视网膜异常与疾病早期发现至关重要。由于其强知识依赖性,眼底图像理解是极具挑战性的多模态任务。现有方法通常依赖内部标注数据进行后训练,但这些数据难以公开,限制了可复现性和研究参与度。为此,我们提出Fundus-R1,一个仅使用公开数据训练的眼科阅读多模态大模型,其中超过94%的数据仅有图像级标签。技术贡献有二:一,提出基于RAG的方法生成与图像相关的、具备眼科知识的推理轨迹;二,通过过程奖励强化推理过程的一致性。在FunBench、Omni-Fundus和GMAI-Fundus三个基准上的实验证明,Fundus-R1显著优于多个基线模型,包括通用模型Qwen2.5-VL及未使用生成推理轨迹的更强版本。该工作为利用公开数据训练强大眼底图像理解模型提供了新路径。

原文摘要 · Abstract (English)

Fundus imaging such as CFP, OCT and UWF is crucial for the early detection of retinal anomalies and diseases. Fundus image understanding, due to its knowledge-intensive nature, poses a challenging vision-language task. An emerging approach to addressing the task is to post-train a generic multimodal large language model (MLLM), either by supervised finetuning (SFT) or by reinforcement learning with verifiable rewards (RLVR), on a considerable amount of in-house samples paired with high-quality clinical reports. However, these valuable samples are not publicly accessible, which not only hinders reproducibility but also practically limits research to few players. To overcome the barrier, we make a novel attempt to train a reasoning-enhanced fundus-reading MLLM, which we term Fundus-R1, using exclusively public datasets, wherein over 94\% of the data are annotated with only image-level labels. Our technical contributions are two-fold. First, we propose a RAG-based method for composing image-specific, knowledge-aware reasoning traces. Such auto-generated traces link visual findings identified by a generic MLLM to the image labels in terms of ophthalmic knowledge. Second, we enhance RLVR with a process reward that encourages self-consistency of the generated reasoning trace in each rollout. Extensive experiments on three fundus-reading benchmarks, i.e., FunBench, Omni-Fundus and GMAI-Fundus, show that Fundus-R1 clearly outperforms multiple baselines, including its generic counterpart (Qwen2.5-VL) and a stronger edition post-trained without using the generated traces. This work paves the way for training powerful fundus-reading MLLMs with publicly available data.

眼科影像多模态推理增强公开数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。