用动态3D高斯建模手术器械,生成逼真合成图像解决数据不足问题。
NeeCo: Image Synthesis of Novel Instrument States Based on Dynamic and Deformable 3D Gaussian Reconstruction
- 基于动态可变形3D高斯点云,实现从新视角和形变下渲染手术器械。
- 合成图像峰值信噪比达29.87,真实感强;训练模型性能提升10%~15%。
- 适合缺乏标注数据的医疗图像研究者,尤其关注手术自动化与数据生成。
基于计算机视觉的技术显著推动了手术自动化的进展,提升了器械追踪、检测与定位能力。然而,现有数据驱动方法依赖大量高质量标注图像数据,限制了其在手术数据科学中的应用。本文提出一种新型动态高斯喷溅技术,以应对手术图像数据稀缺问题。我们构建动态高斯模型来表征动态手术场景,可实现从未见视角与形变下渲染器械,且背景包含真实组织。通过动态训练调整策略应对现实场景中相机位姿校准不佳的问题。此外,提出基于动态高斯的方法,自动为合成数据生成标注。为评估,我们构建了一个新数据集,包含7个场景、14,000帧器械运动与相机位姿变化,以及猪离体组织背景。利用该数据集,可从真实数据中复现场景形变,实现合成图像质量的直接对比。实验表明,本方法生成的图像具有最高峰值信噪比(29.87),达到照片级真实感。进一步在未见过的真实图像数据集上测试基于真实与合成数据训练的医疗专用神经网络,结果显示:使用本方法生成的合成数据训练的模型性能优于当前主流数据增强方法10%,整体模型性能提升近15%。
原文摘要 · Abstract (English)
Computer vision-based technologies significantly enhance surgical automation by advancing tool tracking, detection, and localization. However, Current data-driven approaches are data-voracious, requiring large, high-quality labeled image datasets, which limits their application in surgical data science. Our Work introduces a novel dynamic Gaussian Splatting technique to address the data scarcity in surgical image datasets. We propose a dynamic Gaussian model to represent dynamic surgical scenes, enabling the rendering of surgical instruments from unseen viewpoints and deformations with real tissue backgrounds. We utilize a dynamic training adjustment strategy to address challenges posed by poorly calibrated camera poses from real-world scenarios. Additionally, we propose a method based on dynamic Gaussians for automatically generating annotations for our synthetic data. For evaluation, we constructed a new dataset featuring seven scenes with 14,000 frames of tool and camera motion and tool jaw articulation, with a background of an ex-vivo porcine model. Using this dataset, we synthetically replicate the scene deformation from the ground truth data, allowing direct comparisons of synthetic image quality. Experimental results illustrate that our method generates photo-realistic labeled image datasets with the highest values in Peak-Signal-to-Noise Ratio (29.87). We further evaluate the performance of medical-specific neural networks trained on real and synthetic images using an unseen real-world image dataset. Our results show that the performance of models trained on synthetic images generated by the proposed method outperforms those trained with state-of-the-art standard data augmentation by 10%, leading to an overall improvement in model performances by nearly 15%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。