arXiv:2503.15917cs.CV2025-03被引 2

用少量手术视频高效适配大模型,实现内窥镜3D重建

Learning to Efficiently Adapt Foundation Models for Self-Supervised Endoscopic 3D Scene Reconstruction from Any Cameras

  • 冻结大模型主干,仅训练轻量适配模块提升效率
  • 单网络同时估计深度、位姿与相机参数,精度显著提升
  • 适合医疗视觉研究者,尤其关注少样本场景重建

精准的3D场景重建对多种医疗任务至关重要。由于真实标注数据难以获取,自监督学习(SSL)在内窥镜深度估计中的应用日益增多。尽管基础模型在视觉任务中表现优异,但直接应用于医疗领域常导致性能不佳。然而,其视觉特征仍可增强内窥镜任务,亟需高效的适配策略。本文提出Endo3DAC,一种统一的内窥镜场景重建框架,可高效适配基础模型。设计集成网络,同步估计深度图、相对位姿与相机内参。通过冻结骨干网络,仅训练专用的门控动态向量低秩适配(GDV-LoRA)及独立解码头,实现优越的深度与位姿估计,同时保持训练效率。此外,提出3D重建流水线,优化深度图尺度、偏移及少量参数。在四个内窥镜数据集上的大量实验表明,Endo3DAC显著优于现有方法,且可训练参数更少。据我们所知,这是首个仅需手术视频即可完成自监督深度估计与场景重建的单一网络。代码将在接受后发布。

原文摘要 · Abstract (English)

Accurate 3D scene reconstruction is essential for numerous medical tasks. Given the challenges in obtaining ground truth data, there has been an increasing focus on self-supervised learning (SSL) for endoscopic depth estimation as a basis for scene reconstruction. While foundation models have shown remarkable progress in visual tasks, their direct application to the medical domain often leads to suboptimal results. However, the visual features from these models can still enhance endoscopic tasks, emphasizing the need for efficient adaptation strategies, which still lack exploration currently. In this paper, we introduce Endo3DAC, a unified framework for endoscopic scene reconstruction that efficiently adapts foundation models. We design an integrated network capable of simultaneously estimating depth maps, relative poses, and camera intrinsic parameters. By freezing the backbone foundation model and training only the specially designed Gated Dynamic Vector-Based Low-Rank Adaptation (GDV-LoRA) with separate decoder heads, Endo3DAC achieves superior depth and pose estimation while maintaining training efficiency. Additionally, we propose a 3D scene reconstruction pipeline that optimizes depth maps' scales, shifts, and a few parameters based on our integrated network. Extensive experiments across four endoscopic datasets demonstrate that Endo3DAC significantly outperforms other state-of-the-art methods while requiring fewer trainable parameters. To our knowledge, we are the first to utilize a single network that only requires surgical videos to perform both SSL depth estimation and scene reconstruction tasks. The code will be released upon acceptance.

3D重建自监督医学影像模型适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。