arXiv:2501.05488cs.CVeess.IV2025-01被引 15

基于海量内镜数据训练的通用模型,显著提升胃肠道疾病诊断性能

EndoDINO: A Foundation Model for GI Endoscopy

  • 在10万至千万级内镜图像上预训练ViT模型,构建通用视觉基础模型
  • 仅用简单解码头即在3项任务中达到当前最优表现,泛化能力强
  • 适用于内镜影像分析、医学图像识别等场景,适合临床辅助诊断研究

本文提出EndoDINO,一种面向胃肠道内镜任务的基础模型,通过在文献中已知最大规模的内镜视频数据集中精选的图像数据进行预训练,实现了优异的泛化能力。具体而言,我们使用包含10万至1000万张精心筛选图像的数据集,对参数量分别为10亿、3.07亿和8600万的ViT模型进行了预训练。以冻结的EndoDINO作为特征编码器,仅通过简单的解码头,就在解剖标志分类、息肉分割以及溃疡性结肠炎的梅奥内镜评分(MES)任务中取得了当前最优性能。

原文摘要 · Abstract (English)

In this work, we present EndoDINO, a foundation model for GI endoscopy tasks that achieves strong generalizability by pre-training on a well-curated image dataset sampled from the largest known GI endoscopy video dataset in the literature. Specifically, we pre-trained ViT models with 1B, 307M, and 86M parameters using datasets ranging from 100K to 10M curated images. Using EndoDINO as a frozen feature encoder, we achieved state-of-the-art performance in anatomical landmark classification, polyp segmentation, and Mayo endoscopic scoring (MES) for ulcerative colitis with only simple decoder heads.

内镜分析视觉基础模型医学影像深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。