arXiv:2607.04344cs.CVcs.AI2026-07

IRIS用结构化数据提升眼科疾病视觉问答,实现轻量高效精准诊断。

IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation

论文配图:IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation
图 1 · 摘自论文原文
  • 构建分层主题树与场景驱动对话生成数据,注入临床先验知识。
  • 在12万张眼表图像上达到顶尖性能,超越最大340亿参数模型。
  • 轻量40亿参数模型即可媲美大模型,适合移动端筛查应用。

尽管大型视觉语言模型具备强大通用能力,但在眼表疾病(OSDs)等专业领域,其临床推理受限于高质量多模态指令数据的匮乏。为突破这一数据瓶颈,我们提出IRIS——一种面向眼表疾病细粒度理解的智能识别与交互系统,基于外眼照片进行分析。首先,我们构建了迄今规模最大、最全面的眼表疾病视觉问答数据集IRIS-120K。为克服传统图像-描述对的语义浅层问题,我们提出一种协同数据生成范式,显式注入临床先验知识。数据引擎采用双分支架构:1)主题发现树(TFT),将视觉特征层次化锚定至精确解剖与病理概念,强化医学推理逻辑;2)场景驱动策略,合成角色自适应的临床对话,确保实际泛化能力。通过在该结构化语料库上微调一个紧凑的40亿参数视觉语言模型,IRIS实现顶尖性能,全面优于包括340亿参数在内的通用及专用医疗视觉语言模型。研究结果表明,结构化知识注入远胜于单纯参数扩张,为资源高效、专家级人工智能在移动边缘设备上的可扩展筛查部署开辟可能。代码、数据集与模型权重将公开发布。

原文摘要 · Abstract (English)

While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like ocular surface diseases (OSDs) is severely hindered by a paucity of high-fidelity, multimodal instruction-tuning data. To dismantle this data bottleneck, we introduce IRIS, an Intelligent Recognition and Interaction System tailored for fine-grained OSD understanding via external eye photography. First, we curate IRIS-120K, the largest and most comprehensive OSD visual question-answering (VQA) dataset to date. Crucially, to overcome the semantic shallowness of conventional image-caption pairs, we propose a synergistic data generation paradigm to explicitly inject clinical priors. Our data engine operates via a dual-branch framework: 1) a Topic Finding Tree (TFT) that hierarchically anchors visual features to precise anatomical and pathological concepts, enforcing rigorous medical deduction logic; and 2) a Scene-driven strategy that synthesizes role-adaptive clinical dialogues to ensure pragmatic generalization. By explicitly aligning a compact 4B-parameter VLM on this structurally enriched corpus, IRIS achieves state-of-the-art performance, comprehensively outperforming both generalist and specialized medical VLMs with up to 34B parameters. Our findings underscore that structured knowledge injection profoundly prevails over sheer parameter scaling, unlocking the potential for resource-efficient, expert-level AI deployment on mobile edge devices for scalable OSD screening. Code, datasets, and model weights will be publicly released by this repo.

眼科诊断视觉问答多模态轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。