首个眼科多模态推理数据集与模型,支持从基础到复杂临床思维。
Bridging the Gap in Ophthalmic AI: MM-Retinal-Reason Dataset and OphthaReason Model toward Dynamic Multimodal Reasoning
- 设计动态思维机制,根据样本不确定性调节推理深度。
- 在基础与复杂任务上均超越现有模型至少15%以上。
- 适合医疗AI研究者、眼科诊断系统开发者参考。
多模态大语言模型近期在强化学习框架下展现出强大推理能力。尽管已有若干医学领域多模态推理模型,但大多仅关注基础推理(基于视觉特征匹配的浅层推断)。而真实临床诊断需整合主诉、病史等异构信息与多模态医学影像。为此,我们提出首个涵盖感知与推理全谱的眼科多模态数据集MM-Retinal-Reason,包含基础与复杂推理任务,旨在提升视觉中心推理能力并模拟真实临床思维。基于此数据集,我们构建了首个眼科专用多模态推理模型OphthaReason,具备逐步推理轨迹。为灵活适应不同任务,提出不确定性感知动态思维(UADT)方法,通过熵估计样本级不确定性,并用形变优势机制动态调控探索深度。全面实验表明,该模型在基础与复杂推理任务上均达当前最优,优于通用多模态大模型、医学多模态大模型、基于强化学习的医学多模态大模型及眼科多模态大模型,性能提升分别达24.92%、15.00%、21.20%和17.66%。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have recently demonstrated remarkable reasoning abilities with reinforcement learning paradigm. Although several multimodal reasoning models have been explored in the medical domain, most of them focus exclusively on basic reasoning, which refers to shallow inference based on visual feature matching. However, real-world clinical diagnosis extends beyond basic reasoning, demanding reasoning processes that integrate heterogeneous clinical information (such as chief complaints and medical history) with multimodal medical imaging data. To bridge this gap, we introduce MM-Retinal-Reason, the first ophthalmic multimodal dataset with the full spectrum of perception and reasoning. It encompasses both basic reasoning tasks and complex reasoning tasks, aiming to enhance visual-centric fundamental reasoning capabilities and emulate realistic clinical thinking patterns. Building upon MM-Retinal-Reason, we propose OphthaReason, the first ophthalmology-specific multimodal reasoning model with step-by-step reasoning traces. To enable flexible adaptation to both basic and complex reasoning tasks, we specifically design a novel method called Uncertainty-Aware Dynamic Thinking (UADT), which estimates sample-level uncertainty via entropy and dynamically modulates the model's exploration depth using a shaped advantage mechanism. Comprehensive experiments demonstrate that our model achieves state-of-the-art performance on both basic and complex reasoning tasks, outperforming general-purpose MLLMs, medical MLLMs, RL-based medical MLLMs, and ophthalmic MLLMs by at least 24.92\%, 15.00\%, 21.20\%, and 17.66\%. Project Page: \href{https://github.com/lxirich/OphthaReason}{link}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。