arXiv:2601.22537eess.IVcs.CV2026-01中稿 · publication at IEE…被引 2

端镜图像去模糊去反光并分割,轻量高效适合临床部署

EndoCaver: Handling Fog, Blur and Glare in Endoscopic Images via Joint Deblurring-Segmentation

  • 轻量级双解码器架构,联合优化去模糊与病灶分割
  • 在严重退化图像上仍达0.889 Dice,参数减少90%
  • 适合移动端实时处理,提升内镜筛查可靠性

内镜图像分析对结直肠癌筛查至关重要,但实际场景中常存在镜头起雾、运动模糊和镜面反光,严重影响自动息肉检测。本文提出EndoCaver,一种基于单向引导双解码器的轻量级变压器,实现去模糊与分割的联合多任务处理,显著降低计算复杂度与模型参数。其集成全局注意力模块(GAM)进行跨尺度特征聚合,利用去模糊-分割对齐器(DSA)传递恢复线索,并采用余弦调度器(LoCoS)实现稳定的多任务优化。在Kvasir-SEG数据集上的实验表明,该方法在干净图像上Dice达0.922,在严重退化条件下仍保持0.889,优于现有方法,且模型参数减少90%。结果验证了其高效性与鲁棒性,适用于设备端临床部署。代码已开源。

原文摘要 · Abstract (English)

Endoscopic image analysis is vital for colorectal cancer screening, yet real-world conditions often suffer from lens fogging, motion blur, and specular highlights, which severely compromise automated polyp detection. We propose EndoCaver, a lightweight transformer with a unidirectional-guided dual-decoder architecture, enabling joint multi-task capability for image deblurring and segmentation while significantly reducing computational complexity and model parameters. Specifically, it integrates a Global Attention Module (GAM) for cross-scale aggregation, a Deblurring-Segmentation Aligner (DSA) to transfer restoration cues, and a cosine-based scheduler (LoCoS) for stable multi-task optimisation. Experiments on the Kvasir-SEG dataset show that EndoCaver achieves 0.922 Dice on clean data and 0.889 under severe image degradation, surpassing state-of-the-art methods while reducing model parameters by 90%. These results demonstrate its efficiency and robustness, making it well-suited for on-device clinical deployment. Code is available at https://github.com/ReaganWu/EndoCaver.

医学图像去模糊多任务学习轻量化模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。