arXiv:2605.27916cs.CVcs.CL2026-05被引 1

构建50万条眼科多模态指令数据,提升医疗大模型临床诊断能力。

OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models

论文配图:OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
图 1 · 摘自论文原文
  • 从公开眼科视频中自动提取图文对,构建高质量指令数据
  • 数据含超50万条指令、15万+唯一图像,覆盖问答与推理场景
  • 专为眼科设计,适合医学AI研究者和临床辅助系统开发者

通用医学多模态大模型在构建对话式临床辅助系统方面展现出巨大潜力,但在眼科等高度专业化领域仍缺乏深入探索,主要受限于高质量指令数据的稀缺。现有眼科数据集规模有限,且多依赖公共基准图像,难以捕捉真实临床复杂性。为此,我们提出OphIn-Engine——一种眼科专用指令数据构建流程,从开放获取的眼科网络视频中提取多模态文本与视觉信息。该流程整合了多模态转录、视觉关键线索分离与评分、以及带质量控制的指令生成机制,用于生成准确且多样化的临床对话。基于此,我们构建了包含超过50万条指令实例、151,000余张独特图像及29,000多个视频片段的OphIn-500K数据集,支持视觉问答(VQA)、多轮对话和链式思维(CoT)推理。在此基础上,我们开发了专用于眼科的OphIn-VL多模态大模型,实验证明其性能优于当前最先进的通用与领域特定模型。

原文摘要 · Abstract (English)

The advancement of general medical Multimodal Large Language Models (MLLMs) has shown great potential for building conversational assistants to support clinical diagnosis. However, their adaptation to highly specialized domains such as ophthalmology remains underexplored, primarily due to the scarcity of large-scale, domain-specific instruction-tuning data. Existing ophthalmic datasets for conversational agents are often limited in scale and largely rely on images from established public benchmarks, limiting the scalability of ophthalmic MLLMs and their ability to capture real-world clinical complexity. To address this gap, we propose $\textbf{OphIn-Engine}$, an ophthalmology-specific instruction data curation pipeline that constructs high-quality instruction data from open-access ophthalmology web-scale videos. The pipeline integrates multimodal transcription for extracting image-transcript pairs, visual cue separation and scoring for identifying clinically relevant visual descriptions, and instruction synthesis with quality control for generating accurate and diverse clinical dialogues. Using this engine, we introduce $\textbf{OphIn-500K}$, a large-scale multimodal ophthalmology instruction-tuning dataset containing over 500,000 instruction instances and more than 151,000 unique images from over 29,000 video clips, formatted as visual question answering (VQA), multi-turn conversational interactions, and chain-of-thought (CoT) reasoning. Built upon this dataset, we further develop $\textbf{OphIn-VL}$, an ophthalmology-specific MLLM with advanced visual understanding and conversational capabilities. Comprehensive experiments and case studies demonstrate that OphIn-VL achieves superior performance compared with state-of-the-art general medical and domain-specific MLLMs.

眼科AI多模态指令数据医疗大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。