用专家知识锚定眼科视觉模型,解决误判与细节遗漏问题。
Deep Expert Injection for Anchoring Retinal VLMs with Domain-Specific Knowledge
- 双流编码分离解剖结构与病灶语义,提升细节识别能力。
- 在四个基准上超越大型闭源系统,精准率达新高。
- 适合医疗AI研发者及需要高可信诊断模型的场景。
大型视觉语言模型(LVLM)在自动眼科诊断中潜力巨大,但临床部署受限于缺乏领域专业知识。本文指出两大结构性缺陷:一是感知鸿沟——通用视觉编码器难以捕捉细微病理特征(如微动脉瘤);二是推理鸿沟——深层Transformer层中稀疏视觉证据被海量语言先验覆盖,导致无根据幻觉。为此,提出EyExIn框架,通过深度专家注入机制将专家知识锚定于视网膜LVLM。其采用专家感知双流编码策略,将视觉表征分为通用流(解剖上下文)与专业流(病灶语义)。设计语义自适应门控融合模块,动态增强微小病灶信号并过滤背景噪声。进一步引入自适应深度专家注入,将融合后的视觉特征作为残差偏置嵌入中间LLM层,形成视觉捷径,强制推理过程严格基于视觉证据。在四个基准上的实验表明,该模型持续优于大型专有系统,在眼科视觉问答任务中实现最先进的精确率,推动可信眼科AI的发展。
原文摘要 · Abstract (English)
Large Vision Language Models (LVLMs) show immense potential for automated ophthalmic diagnosis. However, their clinical deployment is severely hindered by lacking domain-specific knowledge. In this work, we identify two structural deficiencies hindering reliable medical reasoning: 1) the Perception Gap, where general-purpose visual encoders fail to resolve fine-grained pathological cues (e.g., microaneurysms); and 2) the Reasoning Gap, where sparse visual evidence is progressively overridden by massive language priors in deeper transformer layers, leading to ungrounded hallucinations. To bridge these gaps, we propose EyExIn, a data-efficient framework designed to anchor retinal VLMs with expert knowledge via a Deep Expert Injection mechanism. Our architecture employs an Expert-Aware Dual-Stream encoding strategy that decouples visual representation into a general stream for anatomical context and a specialized expert stream for pathological semantics. To ensure high-fidelity integration, we design a Semantic-Adaptive Gated Fusion module, which dynamically amplifies subtle lesion signals while filtering irrelevant background noise. Furthermore, we introduce Adaptive Deep Expert Injection to embed persistent "Vision Anchors" by integrating fused visual features as residual biases directly into intermediate LLM layers. This mechanism creates a visual shortcut that forces the reasoning stack to remain strictly grounded in visual evidence. Extensive experiments across four benchmarks demonstrate that our model consistently outperforms massive proprietary systems. EyExIn significantly enhances domain-specific knowledge embedding and achieves state-of-the-art precision in ophthalmic visual question answering, advancing the development of trustworthy ophthalmic AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。