arXiv:2512.23786cs.CVcs.RO2025-12

用合成先验提升手术场景单目深度估计,解决真实手术中反光导致的误差问题。

Bridging the Ex-Vivo to In-Vivo Gap: Synthetic Priors for Monocular Depth Estimation in Specular Surgical Environments

  • 利用高保真合成数据作为先验,通过动态向量低秩适配迁移至医疗领域。
  • 在高反光环境下相对误差降低17%以上,优于现有强基线模型。
  • 首个真实手术数据集ROCALT-90验证,适合临床部署的深度估计研究者参考。

单目深度估计对自主手术机器人至关重要。然而,现有自监督方法普遍存在“离体到活体差距”:在公开数据集上表现良好,但在真实临床应用中性能显著下降。这主要源于实际手术中强烈的镜面反射和液体填充引起的形变。基于噪声真实伪标签训练的模型易出现边界坍缩。为此,本文采用《Depth Anything V2》架构的高保真合成先验,其天然捕捉精确几何细节,并通过动态向量低秩适配(DV-LORA)高效适配至医疗领域。技术上,本方法在公开SCARED数据集上达到新最优;在新型物理分层评估协议下,高反光场景下平方相对误差降低超过17%。此外,为提供领域内严格现实检验,我们引入首个真实手术验证数据集ROCALT-90,包含90个临床腹腔镜序列,具备亚毫米级(<1mm)真实轨迹标注。在该数据集上的评估证明了所提模型在真实临床环境中的优越鲁棒性。

原文摘要 · Abstract (English)

Accurate Monocular Depth Estimation (MDE) is critical for autonomous robotic surgery. However, existing self-supervised methods often exhibit a severe "ex-vivo to in-vivo gap": they achieve high accuracy on public datasets but struggle in actual clinical deployments. This disparity arises because the severe specular reflections and fluid-filled deformations inherent to real surgeries. Models trained on noisy real-world pseudo-labels consequently suffer from severe boundary collapse. To address this, we leverage the high-fidelity synthetic priors of the \textit{Depth Anything V2} architecture, which inherently capture precise geometric details, and efficiently adapt them to the medical domain using Dynamic Vector Low-Rank Adaptation (DV-LORA). Our contributions are two-fold. Technically, our approach establishes a new state-of-the-art on the public SCARED dataset; under a novel physically-stratified evaluation protocol, it reduces Squared Relative Error by over 17\% in high-specularity regimes compared to strong baselines. Furthermore, to provide a rigorous reality check for the field, we introduce \textbf{ROCAL-T 90} (Real Operative CT-Aligned Laparoscopic Trajectories 90), the first real-surgery validation dataset featuring 90 clinical endoscopic sequences with sub-millimeter ($< 1$mm) ground-truth trajectories. Evaluations on ROCAL-T 90 demonstrate our model's superior robustness in true clinical settings.

深度估计手术机器人合成数据医学影像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。