arXiv:2603.25867cs.CV2026-03

用Transformer模型去除手术烟雾,提升内窥镜视觉清晰度

Seeing Through Smoke: Surgical Desmoking for Improved Visual Perception

  • 基于物理启发的Transformer架构,同时预测无烟图像和烟雾分布图
  • 合成8万+对训练数据,真实采集5817对高分辨率图像构建最大烟雾数据集
  • 显著改善图像重建效果,助力深度估计与器械分割,适合医疗影像优化研究者

微创与机器人辅助手术高度依赖内窥镜成像,但电刀和血管闭合设备产生的手术烟雾会严重降低视觉清晰度,阻碍基于视觉的功能。本文提出一种基于Transformer的外科去烟模型,其物理启发式去烟头可联合预测无烟图像与对应烟雾图。为解决成对烟雾到无烟数据稀缺问题,构建了合成数据生成流程,将人工烟雾图案与真实内窥镜图像融合,生成超8万对标注样本用于监督训练。此外,我们首次整理出迄今最大的成对手术烟雾数据集,包含5817对由达芬奇机器人手术系统拍摄的图像,支持高分辨率内窥镜图像的基准测试。在公开基准与自建数据集上的大量实验表明,该方法在图像重建方面优于现有去雾与去烟技术。我们还评估了去烟对下游立体深度估计与器械分割的影响,揭示数字去烟方法的潜力与当前局限。

原文摘要 · Abstract (English)

Minimally invasive and robot-assisted surgery relies heavily on endoscopic imaging, yet surgical smoke produced by electrocautery and vessel-sealing instruments can severely degrade visual perception and hinder vision-based functionalities. We present a transformer-based surgical desmoking model with a physics-inspired desmoking head that jointly predicts smoke-free image and corresponding smoke map. To address the scarcity of paired smoky-to-smoke-free training data, we develop a synthetic data generation pipeline that blends artificial smoke patterns with real endoscopic images, yielding over 80,000 paired samples for supervised training. We further curate, to our knowledge, the largest paired surgical smoke dataset to date, comprising 5,817 image pairs captured with the da Vinci robotic surgical system, enabling benchmarking on high-resolution endoscopic images. Extensive experiments on both a public benchmark and our dataset demonstrate state-of-the-art performance in image reconstruction compared to existing dehazing and desmoking approaches. We also assess the impact of desmoking on downstream stereo depth estimation and instrument segmentation, highlighting both the potential benefits and current limitations of digital smoke removal methods.

医学图像去烟Transformer机器人手术

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。