提出新任务与框架,让模型像医生一样分析超长肠镜视频并生成可靠诊断。
Divide-then-Diagnose: Weaving Clinician-Inspired Contexts for Ultra-Long Capsule Endoscopy Videos

- 模仿医生读片流程,分步筛选关键帧并构建诊断上下文。
- 在240段真实肠镜视频上实现比现有方法更准确的诊断摘要。
- 适合医疗影像分析、临床辅助诊断研究者参考。
胶囊内镜(CE)可实现无创胃肠筛查,但现有研究多局限于帧级分类与检测,视频级分析仍不充分。为此,我们提出并定义了一项新任务:基于诊断的胶囊内镜视频摘要,要求从海量视频中提取包含临床意义发现的关键帧,并据此做出准确诊断。该任务极具挑战性:诊断相关事件极其稀疏,易被数万张正常帧淹没;且单帧常因运动模糊、污物、反光和视角快速变化而模糊不清。为推动该方向研究,我们构建了首个带有诊断驱动标注的视频数据集VideoCAP,包含240段完整视频,提供真实临床报告支持的关键帧提取与诊断标注。针对此任务,我们进一步提出DiCE框架,其灵感源自临床读片流程:先对原始视频进行高效候选帧筛选,再通过上下文编织器将候选帧组织成连贯的诊断上下文以保留不同病变事件,最后由证据汇聚器在每个上下文中聚合多帧证据,形成鲁棒的片段级判断。实验表明,DiCE持续优于现有先进方法,生成简洁且临床可信的诊断摘要。结果凸显基于诊断的上下文推理是超长胶囊内镜视频摘要的有前景范式。
原文摘要 · Abstract (English)
Capsule endoscopy (CE) enables non-invasive gastrointestinal screening, but current CE research remains largely limited to frame-level classification and detection, leaving video-level analysis underexplored. To bridge this gap, we introduce and formally define a new task, diagnosis-driven CE video summarization, which requires extracting key evidence frames that covers clinically meaningful findings and making accurate diagnoses from those evidence frames. This setting is challenging because diagnostically relevant events are extremely sparse and can be overwhelmed by tens of thousands of redundant normal frames, while individual observations are often ambiguous due to motion blur, debris, specular highlights, and rapid viewpoint changes. To facilitate research in this direction, we introduce VideoCAP, the first CE dataset with diagnosis-driven annotations derived from real clinical reports. VideoCAP comprises 240 full-length videos and provides realistic supervision for both key evidence frame extraction and diagnosis. To address this task, we further propose DiCE, a clinician-inspired framework that mirrors the standard CE reading workflow. DiCE first performs efficient candidate screening over the raw video, then uses a Context Weaver to organize candidates into coherent diagnostic contexts that preserve distinct lesion events, and an Evidence Converger to aggregate multi-frame evidence within each context into robust clip-level judgments. Experiments show that DiCE consistently outperforms state-of-the-art methods, producing concise and clinically reliable diagnostic summaries. These results highlight diagnosis-driven contextual reasoning as a promising paradigm for ultra-long CE video summarization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。