arXiv:2601.18929cs.CV2026-01

用深度信息预训练,让手术视觉模型更高效

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training

  • 在140万组手术图像上融合深度图进行多模态预训练
  • 仅用25%标注数据就超过全量数据的单模态模型
  • 无需改动推理结构,深度仅用于预训练,易落地

视觉基础模型(VFMs)已成为手术场景理解的强大工具。但现有方法多依赖单模态RGB预训练,忽略了手术环境中的复杂三维几何特征。尽管已有架构支持多模态或几何感知输入,但在手术场景中引入深度信息的优势仍缺乏研究。我们开展大规模实证研究,比较八种基于ViT的VFMs,其差异在于预训练领域、学习目标和输入模态(RGB vs. RGB-D)。预训练使用包含140万组机器人手术图像及其由现成网络生成的深度图的精选数据集。在八个涵盖目标检测、分割、深度估计和位姿估计的手术数据集上,评估了冻结主干和端到端微调两种策略。实验结果一致显示:采用显式几何标记化的模型(如MultiMAE)在所有任务中显著优于单模态基线。值得注意的是,几何感知预训练带来显著的数据效率提升:仅用25%标注数据微调的模型,始终优于使用全量数据训练的仅RGB模型。重要的是,这些收益无需在推理时进行架构或运行时修改;深度信息仅在预训练阶段使用,部署极为简便。结果表明,多模态预训练为构建更强大的手术视觉系统提供了可行路径。

原文摘要 · Abstract (English)

Vision foundation models (VFMs) have emerged as powerful tools for surgical scene understanding. However, current approaches predominantly rely on unimodal RGB pre-training, overlooking the complex 3D geometry inherent to surgical environments. Although several architectures support multimodal or geometry-aware inputs in general computer vision, the benefits of incorporating depth information in surgical settings remain underexplored. We conduct a large-scale empirical study comparing eight ViT-based VFMs that differ in pre-training domain, learning objective, and input modality (RGB vs. RGB-D). For pre-training, we use a curated dataset of 1.4 million robotic surgical images paired with depth maps generated from an off-the-shelf network. We evaluate these models under both frozen-backbone and end-to-end fine-tuning protocols across eight surgical datasets spanning object detection, segmentation, depth estimation, and pose estimation. Our experiments yield several consistent findings. Models incorporating explicit geometric tokenization, such as MultiMAE, substantially outperform unimodal baselines across all tasks. Notably, geometric-aware pre-training enables remarkable data efficiency: models fine-tuned on just 25% of labeled data consistently surpass RGB-only models trained on the full dataset. Importantly, these gains require no architectural or runtime changes at inference; depth is used only during pre-training, making adoption straightforward. These findings suggest that multimodal pre-training offers a viable path towards building more capable surgical vision systems.

手术视觉多模态深度预训练数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。