用双向注意力机制提升手术烟雾下的视觉清晰度,助力机器人手术精准操作。
RGA-Net: A Vision Enhancement Framework for Robotic Surgical Systems Using Reciprocal Attention Mechanisms
- 通过双流混合注意力与轴分解注意力,协同捕捉局部细节与全局光照变化。
- 在DesmokeData和LSD3K数据集上显著提升图像恢复质量,实现稳定清晰可视化。
- 适合需提升手术视觉可靠性的人机协作系统研发者及临床工程团队。
机器人手术系统高度依赖高质量视觉反馈以实现精确远程操作;然而,能量设备产生的手术烟雾会严重降低内窥镜视频画质,影响人机交互与手术效果。本文提出RGA-Net(双向门控与注意力融合网络),一种专为机器人手术流程中烟雾去除设计的深度学习框架。针对手术烟雾具有密度高、分布不均、光散射复杂等特性,采用分层编码器-解码器结构,包含两项核心创新:(1) 双流混合注意力(DHA)模块,结合移位窗口注意力与频域处理,同时捕捉局部手术细节与全局光照变化;(2) 轴分解注意力(ADA)模块,通过因子化注意力机制高效处理多尺度特征。编码器与解码器路径间通过双向交叉门控块实现特征双向调制。在DesmokeData与LSD3K手术数据集上的大量实验表明,RGA-Net在恢复视觉清晰度方面表现优异,适用于机器人手术集成。该方法通过提供持续清晰的可视化,改善人机交互,减轻术者认知负担,优化操作流程,降低医源性损伤风险,为未来临床试验中的外科医生可用性评估奠定基础。该框架代表了通过计算视觉增强提升机器人手术系统可靠性与安全性的关键进展。
原文摘要 · Abstract (English)
Robotic surgical systems rely heavily on high-quality visual feedback for precise teleoperation; yet, surgical smoke from energy-based devices significantly degrades endoscopic video feeds, compromising the human-robot interface and surgical outcomes. This paper presents RGA-Net (Reciprocal Gating and Attention-fusion Network), a novel deep learning framework specifically designed for smoke removal in robotic surgery workflows. Our approach addresses the unique challenges of surgical smoke-including dense, non-homogeneous distribution and complex light scattering-through a hierarchical encoder-decoder architecture featuring two key innovations: (1) a Dual-Stream Hybrid Attention (DHA) module that combines shifted window attention with frequency-domain processing to capture both local surgical details and global illumination changes, and (2) an Axis-Decomposed Attention (ADA) module that efficiently processes multi-scale features through factorized attention mechanisms. These components are connected via reciprocal cross-gating blocks that enable bidirectional feature modulation between encoder and decoder pathways. Extensive experiments on the DesmokeData and LSD3K surgical datasets demonstrate that RGA-Net achieves superior performance in restoring visual clarity suitable for robotic surgery integration. Our method enhances the surgeon-robot interface by providing consistently clear visualization, laying a technical foundation for alleviating surgeons' cognitive burden, optimizing operation workflows, and reducing iatrogenic injury risks in minimally invasive procedures. These practical benefits could be further validated through future clinical trials involving surgeon usability assessments. The proposed framework represents a significant step toward more reliable and safer robotic surgical systems through computational vision enhancement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。