首个多任务动态手术场景重建与分割系统,实时高效且精度领先。
Feature-EndoGaussian: Feature Distilled Gaussian Splatting in Surgical Deformable Scene Reconstruction
- 用预训练语义嵌入蒸馏特征,实现可变形场景的统一4D表示。
- 61帧/秒实时渲染,在EndoNeRF上达39.1 PSNR,SCARED上27.3 PSNR。
- 适合需要高保真实时视觉反馈的微创手术辅助系统研发者。
微创手术需要对动态、低纹理的手术场景提供高保真、实时的视觉反馈。为此,我们提出首个基于特征蒸馏4D高斯点阵的实时管道Feature-EndoGaussian(FE-4DGS),实现可变形手术环境的同时重建与语义分割。不同于以往仅限静态场景的特征蒸馏方法,也区别于缺乏语义融合的现有4D方法,FE-4DGS通过预训练2D语义嵌入,生成一个统一的4D表示——语义信息随组织运动同步变形。该统一方法通过单一并行化光栅化过程,实现实时RGB与语义输出。尽管引入了特征蒸馏的额外复杂度,FE-4DGS仍保持61 FPS实时渲染,模型紧凑;在EndoNeRF数据集上达到39.1 PSNR的顶尖渲染保真度,在SCARED上达27.3 PSNR;语义分割性能在EndoVis18上表现优异,二分类任务达0.93 DSC,多标签任务达0.77 DSC,媲美甚至超越强基线2D模型。
原文摘要 · Abstract (English)
Minimally invasive surgery (MIS) requires high-fidelity, real-time visual feedback of dynamic and low-texture surgical scenes. To address these requirements, we introduce FeatureEndo-4DGS (FE-4DGS), the first real time pipeline leveraging feature-distilled 4D Gaussian Splatting for simultaneous reconstruction and semantic segmentation of deformable surgical environments. Unlike prior feature-distilled methods restricted to static scenes, and existing 4D approaches that lack semantic integration, FE-4DGS seamlessly leverages pre-trained 2D semantic embeddings to produce a unified 4D representation-where semantics also deform with tissue motion. This unified approach enables the generation of real-time RGB and semantic outputs through a single, parallelized rasterization process. Despite the additional complexity from feature distillation, FE-4DGS sustains real-time rendering (61 FPS) with a compact footprint, achieves state-of-the-art rendering fidelity on EndoNeRF (39.1 PSNR) and SCARED (27.3 PSNR), and delivers competitive EndoVis18 segmentation, matching or exceeding strong 2D baselines for binary segmentation tasks (0.93 DSC) and remaining competitive for multi-label segmentation (0.77 DSC).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。