构建首个内窥镜图像修复基准数据集,解决术中烟雾、雾气和污渍问题。
Benchmarking Endoscopic Surgical Image Restoration and Beyond
- 构建包含3113张图像的SurgClean数据集,覆盖去烟、去雾、去溅污三类任务。
- 22种算法在该数据集上表现均未达临床要求,暴露算法提升空间。
- 揭示医源与自然场景在结构感知和语义理解上的差异,助力专用修复研究。
内窥镜手术中清晰高质量的视觉对医生做出准确术中决策至关重要。然而,能量设备产生的烟雾、温差导致的镜头起雾、血液或组织液飞溅造成的镜头污染等持续性视觉退化,严重损害图像清晰度,干扰手术流程并威胁患者安全。为系统研究和解决各类手术场景退化问题,本文提出一个真实世界开源的内窥镜图像修复数据集SurgClean,涵盖两个医疗场景下的多类型图像修复任务:去烟、去雾和去溅污。SurgClean包含3,113张具有多样退化类型及对应配对参考标签的图像。基于此数据集,我们建立标准化评估基准,并测试了22种代表性通用及任务特定图像修复方法(12种通用+10种专用)。实验结果表明,现有方法性能与临床需求仍有显著差距,凸显智能手术图像修复算法发展的关键机遇。此外,我们从结构感知与语义理解角度分析医源与自然场景退化差异,为领域特定图像修复研究提供基础洞见。本工作旨在推动修复算法发展,提升临床操作效率。
原文摘要 · Abstract (English)
In endoscopic surgery, a clear and high-quality visual field is critical for surgeons to make accurate intraoperative decisions. However, persistent visual degradation, including smoke generated by energy devices, lens fogging from thermal gradients, and lens contamination due to blood or tissue fluid splashes during surgical procedures, severely impairs visual clarity. These degenerations can seriously hinder surgical workflow and pose risks to patient safety. To systematically investigate and address various forms of surgical scene degradation, we introduce a real- world open-source surgical image restoration dataset covering endoscopic environments, called SurgClean, which involves multi-type image restoration tasks from two medical sites, i.e., desmoking, defogging, and desplashing. SurgClean comprises 3,113 images with diverse degradation types and corresponding paired reference labels. Based on SurgClean, we establish a standardized evaluation benchmark and provide performance for 22 representative generic task-specific image restoration approaches, including 12 generic and 10 task-specific image restoration approaches. Experimental results reveal substantial performance gaps relative to clinical requirements, highlighting a critical opportunity for algorithm advancements in intelligent surgical restoration. Furthermore, we explore the degradation discrepancies between surgical and natural scenes from structural perception and semantic under- standing perspectives, providing fundamental insights for domain-specific image restoration research. Our work aims to empower restoration algorithms and improve the efficiency of clinical procedures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。