arXiv:2502.20669cs.CV2025-02被引 2

通过物理渲染分离材质与光照,实现手术场景的逼真新视角合成。

EndoPBR: Material and Lighting Estimation for Photorealistic Surgical Simulations via Physically-based Rendering

  • 显式解耦材质与光照,基于渲染方程生成真实图像。
  • 在结肠镜视频数据集上达到与现有方法相当的新视角合成效果。
  • 生成的合成数据可有效微调深度估计模型,媲美真实数据训练。

外科场景中缺乏标注的3D视觉数据集制约了鲁棒3D重建算法的发展。尽管神经辐射场和3D高斯泼溅在通用计算机视觉领域流行,但因非静态光照和非朗伯表面等挑战,其在手术场景中尚未取得稳定成功。为此,本文提出一种从内窥镜图像与已知几何结构中进行材质与光照估计的可微分渲染框架。不同于以往将光照与材质联合建模为辐射率的方法,我们显式解耦二者以实现更稳健、逼真的新视角合成。为消除训练歧义,我们引入外科场景特有的先验:将光照建模为简单聚光灯,材质则采用由神经网络参数化的双向反射分布函数。通过将颜色预测建立在渲染方程基础上,可在任意相机姿态下生成逼真图像。我们在结肠镜3D视频数据集上的多个序列上进行了评估,结果表明该方法在新视角合成方面表现优异。此外,我们还证明合成数据可用于开发3D视觉算法——通过使用生成结果微调深度估计模型,其性能与使用原始真实图像微调相当。

原文摘要 · Abstract (English)

The lack of labeled datasets in 3D vision for surgical scenes inhibits the development of robust 3D reconstruction algorithms in the medical domain. Despite the popularity of Neural Radiance Fields and 3D Gaussian Splatting in the general computer vision community, these systems have yet to find consistent success in surgical scenes due to challenges such as non-stationary lighting and non-Lambertian surfaces. As a result, the need for labeled surgical datasets continues to grow. In this work, we introduce a differentiable rendering framework for material and lighting estimation from endoscopic images and known geometry. Compared to previous approaches that model lighting and material jointly as radiance, we explicitly disentangle these scene properties for robust and photorealistic novel view synthesis. To disambiguate the training process, we formulate domain-specific properties inherent in surgical scenes. Specifically, we model the scene lighting as a simple spotlight and material properties as a bidirectional reflectance distribution function, parameterized by a neural network. By grounding color predictions in the rendering equation, we can generate photorealistic images at arbitrary camera poses. We evaluate our method with various sequences from the Colonoscopy 3D Video Dataset and show that our method produces competitive novel view synthesis results compared with other approaches. Furthermore, we demonstrate that synthetic data can be used to develop 3D vision algorithms by finetuning a depth estimation model with our rendered outputs. Overall, we see that the depth estimation performance is on par with fine-tuning with the original real images.

医学图像物理渲染新视角合成3D重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。