针对肠镜视频中运动模糊和尺度变化,提出自适应检测框架提升准确率。
AVPDN: Learning Motion-Robust and Scale-Adaptive Representations for Video-Based Polyp Detection
- 设计双模块结构:自适应特征交互与多尺度上下文融合。
- 在多个公开数据集上达到领先性能,显著降低误检率。
- 适合需要高鲁棒性视频病理检测的临床辅助系统开发。
准确检测息肉对结直肠癌早期和中期诊断至关重要。与静态图像相比,动态肠镜视频提供更全面的视觉信息,有助于制定有效治疗方案。然而,不同于固定摄像头录制,肠镜视频常伴随快速镜头移动,引入大量背景噪声,破坏场景结构完整性,增加误报风险。为此,本文提出自适应视频息肉检测网络(AVPDN),一种针对肠镜视频多尺度息肉检测的鲁棒框架。AVPDN包含两个核心组件:自适应特征交互与增强(AFIA)模块和感知尺度的上下文融合(SACI)模块。AFIA采用三分支架构增强特征表示,结合密集自注意力进行全局上下文建模,稀疏自注意力缓解特征聚合中查询-键相似度低的问题,并通过通道混洗促进跨分支信息交换。同时,SACI模块利用不同感受野的空洞卷积捕捉多空间尺度上下文信息,提升模型去噪能力。在多个具有挑战性的公开基准上的实验表明,所提方法在视频息肉检测任务中表现优异且具备良好泛化能力。
原文摘要 · Abstract (English)
Accurate detection of polyps is of critical importance for the early and intermediate stages of colorectal cancer diagnosis. Compared to static images, dynamic colonoscopy videos provide more comprehensive visual information, which can facilitate the development of effective treatment plans. However, unlike fixed-camera recordings, colonoscopy videos often exhibit rapid camera movement, introducing substantial background noise that disrupts the structural integrity of the scene and increases the risk of false positives. To address these challenges, we propose the Adaptive Video Polyp Detection Network (AVPDN), a robust framework for multi-scale polyp detection in colonoscopy videos. AVPDN incorporates two key components: the Adaptive Feature Interaction and Augmentation (AFIA) module and the Scale-Aware Context Integration (SACI) module. The AFIA module adopts a triple-branch architecture to enhance feature representation. It employs dense self-attention for global context modeling, sparse self-attention to mitigate the influence of low query-key similarity in feature aggregation, and channel shuffle operations to facilitate inter-branch information exchange. In parallel, the SACI module is designed to strengthen multi-scale feature integration. It utilizes dilated convolutions with varying receptive fields to capture contextual information at multiple spatial scales, thereby improving the model's denoising capability. Experiments conducted on several challenging public benchmarks demonstrate the effectiveness and generalization ability of the proposed method, achieving competitive performance in video-based polyp detection tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。