arXiv:2602.22821cs.CV2026-02被引 1

通过时序因果聚合与自适应参考帧,提升视频肠镜息肉分割精度与实时性

CMSA-Net: Causal Multi-scale Aggregation with Adaptive Multi-source Reference for Video Polyp Segmentation

  • 引入因果多尺度聚合模块,按时间顺序融合多帧语义信息
  • 在SUN-SEG数据集上达到最新最优性能,兼顾准确率与实时推理
  • 自适应选择可靠参考帧,适合临床实时辅助诊断场景

视频息肉分割(VPS)是计算机辅助结肠镜检查中的关键任务,有助于医生精准定位和追踪息肉。然而,由于息肉与周围黏膜外观相似,导致语义区分能力弱;且各帧间息肉位置与尺度变化大,难以实现稳定准确的分割。为此,本文提出稳健的VPS框架CMSA-Net。网络引入因果多尺度聚合(CMA)模块,从多历史帧中按不同尺度有效聚合语义信息。通过因果注意力机制,确保特征传播严格遵循时间顺序,减少噪声并提升特征可靠性。此外,设计动态多源参考(DMR)策略,基于语义可分性和预测置信度自适应选择高信息量、可靠的参考帧,提供强多帧指导的同时保持模型高效,适用于实时推理。在SUN-SEG数据集上的大量实验表明,CMSA-Net取得当前最优性能,实现了分割精度与临床实时应用性的良好平衡。

原文摘要 · Abstract (English)

Video polyp segmentation (VPS) is an important task in computer-aided colonoscopy, as it helps doctors accurately locate and track polyps during examinations. However, VPS remains challenging because polyps often look similar to surrounding mucosa, leading to weak semantic discrimination. In addition, large changes in polyp position and scale across video frames make stable and accurate segmentation difficult. To address these challenges, we propose a robust VPS framework named CMSA-Net. The proposed network introduces a Causal Multi-scale Aggregation (CMA) module to effectively gather semantic information from multiple historical frames at different scales. By using causal attention, CMA ensures that temporal feature propagation follows strict time order, which helps reduce noise and improve feature reliability. Furthermore, we design a Dynamic Multi-source Reference (DMR) strategy that adaptively selects informative and reliable reference frames based on semantic separability and prediction confidence. This strategy provides strong multi-frame guidance while keeping the model efficient for real-time inference. Extensive experiments on the SUN-SEG dataset demonstrate that CMSA-Net achieves state-of-the-art performance, offering a favorable balance between segmentation accuracy and real-time clinical applicability.

视频分割医学图像因果建模实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。