arXiv:2506.15200cs.CV2025-06被引 3

让眼科OCT图像模型通过少量样本实时学习新任务,突破传统模型局限。

Conquering the Retina: Bringing Visual in-Context Learning to OCT

  • 用视觉上下文学习让模型在推理时根据少量示例快速适应新任务
  • 在多个OCT数据集上建立首个基准,验证该方法的潜力与不足
  • 开源代码推动医疗AI通用模型发展,适合临床研究者使用

医学影像分析近年发展出针对特定临床任务的高度专业化模型,性能优异但仅限预设任务,需专家和大量资源开发与适配。相比之下,通用模型可让医生即时定义新任务而无需专门建模。本文探索在视网膜OCT领域应用视觉上下文学习(VICL),即训练模型基于推理时提供的少量示例跨任务泛化。为实现严格评估,提出专用于OCT的VICL评价协议,对当前最先进的医疗VICL方法在多个视网膜OCT数据集上进行广泛测试,建立首个基线,揭示该技术在OCT中的潜力与现存局限。为促进后续研究与实际应用,公开发布代码。

原文摘要 · Abstract (English)

Recent advancements in medical image analysis have led to the development of highly specialized models tailored to specific clinical tasks. These models have demonstrated exceptional performance and remain a crucial research direction. Yet, their applicability is limited to predefined tasks, requiring expertise and extensive resources for development and adaptation. In contrast, generalist models offer a different form of utility: allowing medical practitioners to define tasks on the fly without the need for task-specific model development. In this work, we explore how to train generalist models for the domain of retinal optical coherence tomography using visual in-context learning (VICL), i.e., training models to generalize across tasks based on a few examples provided at inference time. To facilitate rigorous assessment, we propose a broad evaluation protocol tailored to VICL in OCT. We extensively evaluate a state-of-the-art medical VICL approach on multiple retinal OCT datasets, establishing a first baseline to highlight the potential and current limitations of in-context learning for OCT. To foster further research and practical adoption, we openly release our code.

医学影像上下文学习OCT通用模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。