用大模型引导的点阵技术,实现内窥镜场景的物理仿真
EndoGSim: Physics-Aware 4D Dynamic Endoscopic Scene Simulations via MLLM-Guided Gaussian Splatting

- 用多模态大模型指导点阵法,重建可变形组织与器械
- 自动生成材料属性,通过物理引擎优化,提升仿真真实性
- 适合机器人手术模拟与训练系统开发者使用
在机器人辅助微创手术中,高保真动态内窥镜场景重建与仿真对提升下游任务和手术效果至关重要。然而,现有方法主要关注视觉重建,缺乏场景的物理描述以支持真实仿真。本文提出统一框架,通过多模态大语言模型(MLLM)引导的4D高斯点阵(4DGS),实现内窥镜场景的物理感知重建与物理仿真。方法结合预训练分割与深度估计,用4DGS表示可变形组织与器械。为自动推断物理属性,引入基于物体的材料场,由MLLM初始化材料参数,并通过可微分物质点法(MPM)在图像与光流联合监督下优化。在开源与自研数据集上验证,本框架在仿真保真度与物理准确性上均优于当前最优方法,展现出推动机器人辅助手术应用的巨大潜力。
原文摘要 · Abstract (English)
In robot-assisted minimally invasive surgery, high-fidelity dynamic endoscopic scene reconstruction and simulation are crucial to enhancing downstream tasks and advancing surgical outcomes. However, existing methods primarily focus on visual reconstruction, lacking physics-based descriptions of the scene required for realistic simulation. We propose a unified framework that achieves physics-aware reconstruction and physical simulation of endoscopic scenes through Multi-modal Large Language Models (MLLMs)-guided Gaussian Splatting. Our approach utilizes 4D Gaussian Splatting (4DGS) integrated with pre-trained segmentation and depth estimation to represent deformable tissues and tools. To achieve automatic inference of physical properties, we introduce an object-wise material field that initializes material parameters via MLLM and refines them through a differentiable Material Point Method (MPM) under joint supervision from rendered images and optical flow. Validated on both open-source and in-house datasets, our framework achieves superior simulation fidelity and physical accuracy compared to state-of-the-art methods, underscoring its potential to advance robot-assisted surgical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。