arXiv:2603.25645eess.IVcs.CV2026-03被引 1

构建首个全流程结肠镜视频密集标注数据集,助力医疗大模型评估与优化。

Colon-Bench: An Agentic Workflow for Scalable Dense Lesion Annotation in Full-Procedure Colonoscopy Videos

  • 通过多阶段智能工作流实现长视频的高效标注
  • 覆盖528段视频、超30万框、13万字临床描述
  • 提出新提示策略提升医疗大模型零样本表现

早期通过结肠镜筛查对预防结直肠癌至关重要,但开发鲁棒的AI系统受限于缺乏密集标注的长序列视频数据集。现有数据集主要聚焦单类息肉检测,缺乏评估现代多模态大语言模型(MLLMs)所需的丰富空间、时间与语言标注。为此,我们提出Colon-Bench,采用创新的多阶段智能工作流构建。该流程融合时序提案、边界框追踪、AI驱动的视觉确认及人工审核,可规模化标注完整手术过程视频。最终验证的数据集规模空前,包含528段视频、14种病变类别(包括息肉、溃疡、出血等)、超过30万个人工标注框、21.3万张分割掩码以及13.3万词临床描述。我们利用Colon-Bench严格评估了前沿MLLM在病变分类、开放词汇视频目标分割(OV-VOS)和视频视觉问答(VQA)任务上的表现。结果显示,相比SAM-3,MLLM在医学领域展现出惊人的定位能力。进一步分析发现常见VQA错误,并提出新颖的“结肠技能”提示策略,使多数MLLM的零样本性能提升最高达9.7%。数据集与代码已公开:https://abdullahamdi.com/colon-bench。

原文摘要 · Abstract (English)

Early screening via colonoscopy is critical for colon cancer prevention, yet developing robust AI systems for this domain is hindered by the lack of densely annotated, long-sequence video datasets. Existing datasets predominantly focus on single-class polyp detection and lack the rich spatial, temporal, and linguistic annotations required to evaluate modern Multimodal Large Language Models (MLLMs). To address this critical gap, we introduce Colon-Bench, generated via a novel multi-stage agentic workflow. Our pipeline seamlessly integrates temporal proposals, bounding-box tracking, AI-driven visual confirmation, and human-in-the-loop review to scalably annotate full-procedure videos. The resulting verified benchmark is unprecedented in scope, encompassing 528 videos, 14 distinct lesion categories (including polyps, ulcers, and bleeding), over 300,000 bounding boxes, 213,000 segmentation masks, and 133,000 words of clinical descriptions. We utilize Colon-Bench to rigorously evaluate state-of-the-art MLLMs across lesion classification, Open-Vocabulary Video Object Segmentation (OV-VOS), and video Visual Question Answering (VQA). The MLLM results demonstrate surprisingly high localization performance in medical domains compared to SAM-3. Finally, we analyze common VQA errors from MLLMs to introduce a novel "colon-skill" prompting strategy, improving zero-shot MLLM performance by up to 9.7% across most MLLMs. The dataset and the code are available at https://abdullahamdi.com/colon-bench .

医学影像视频标注大模型评估结肠镜

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。