arXiv:2503.24108cs.CVcs.AI2025-03被引 5

统一模型同时完成肠镜视频中的息肉检测、分割、分类与追踪,无需微调。

PolypSegTrack: Unified Foundation Model for Colonoscopy Video Analysis

  • 用条件掩码损失实现跨数据集灵活训练,支持像素级或框标注
  • 无监督追踪模块通过对象查询关联帧间息肉,不依赖启发式规则
  • 基于自然图像预训练的视觉主干,避免领域特异性预训练

结肠镜检查中早期发现、准确分割、分类和追踪息肉对预防结直肠癌至关重要。现有深度学习方法要么需要针对任务微调,缺乏追踪能力,或依赖领域特定预训练。本文提出PolypSegTrack,一种新型基础模型,可联合解决结肠镜视频中的息肉检测、分割、分类与无监督追踪。该方法引入新型条件掩码损失,使模型能灵活训练于具有像素级分割标签或边界框标注的数据集,从而避免任务特异性微调。无监督追踪模块利用对象查询可靠关联多帧间的息肉实例,不依赖任何启发式规则。我们采用在自然图像上无监督预训练的稳健视觉主干,消除领域特异性预训练需求。在多个息肉基准上的大量实验表明,该方法在检测、分割、分类和追踪任务上均显著优于现有最先进方法。

原文摘要 · Abstract (English)

Early detection, accurate segmentation, classification and tracking of polyps during colonoscopy are critical for preventing colorectal cancer. Many existing deep-learning-based methods for analyzing colonoscopic videos either require task-specific fine-tuning, lack tracking capabilities, or rely on domain-specific pre-training. In this paper, we introduce PolypSegTrack, a novel foundation model that jointly addresses polyp detection, segmentation, classification and unsupervised tracking in colonoscopic videos. Our approach leverages a novel conditional mask loss, enabling flexible training across datasets with either pixel-level segmentation masks or bounding box annotations, allowing us to bypass task-specific fine-tuning. Our unsupervised tracking module reliably associates polyp instances across frames using object queries, without relying on any heuristics. We leverage a robust vision foundation model backbone that is pre-trained unsupervisedly on natural images, thereby removing the need for domain-specific pre-training. Extensive experiments on multiple polyp benchmarks demonstrate that our method significantly outperforms existing state-of-the-art approaches in detection, segmentation, classification, and tracking.

医学影像目标追踪基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。