用动态图网络提升手术中软组织3D重建精度与泛化能力
EndoVGGT: GNN-Enhanced Depth Estimation for Surgical 3D Reconstruction
- 通过动态构建语义图捕捉组织间长程关联,突破静态邻域限制
- 在SCARED数据集上PSNR提升24.6%,SSIM提升9.1%,显著改善重建质量
- 零样本跨数据集泛化能力强,适合需高鲁棒性的外科视觉系统
准确重建可变形软组织的3D结构对术中机器人感知至关重要。然而,低纹理表面、反光和器械遮挡常破坏几何连续性,给现有固定拓扑方法带来挑战。为此,我们提出EndoVGGT,一种以几何为中心的框架,配备变形感知图注意力(DeGAT)模块。DeGAT不依赖静态空间邻域,而是动态构建特征空间语义图,捕捉相干组织区域间的长程相关性,实现遮挡区域间结构线索的鲁棒传播,强化全局一致性并提升非刚性形变恢复能力。在SCARED数据集上的大量实验表明,本方法显著提升重建保真度,相比先前最先进方法,PSNR提高24.6%,SSIM提升9.1%。关键的是,EndoVGGT在未见的SCARED与EndoNeRF数据集间展现出强大的零样本跨数据集泛化能力,证实DeGAT学习到了领域无关的几何先验。这些结果凸显了动态特征空间建模在一致手术3D重建中的有效性。
原文摘要 · Abstract (English)
Accurate 3D reconstruction of deformable soft tissues is essential for surgical robotic perception. However, low-texture surfaces, specular highlights, and instrument occlusions often fragment geometric continuity, posing a challenge for existing fixed-topology approaches. To address this, we propose EndoVGGT, a geometry-centric framework equipped with a Deformation-aware Graph Attention (DeGAT) module. Rather than using static spatial neighborhoods, DeGAT dynamically constructs feature-space semantic graphs to capture long-range correlations among coherent tissue regions. This enables robust propagation of structural cues across occlusions, enforcing global consistency and improving non-rigid deformation recovery. Extensive experiments on SCARED show that our method significantly improves fidelity, increasing PSNR by 24.6% and SSIM by 9.1% over prior state-of-the-art. Crucially, EndoVGGT exhibits strong zero-shot cross-dataset generalization to the unseen SCARED and EndoNeRF domains, confirming that DeGAT learns domain-agnostic geometric priors. These results highlight the efficacy of dynamic feature-space modeling for consistent surgical 3D reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。