arXiv:2410.00979cs.CVcs.AI2024-10

提出全参数高效自学习框架,提升内窥镜深度估计精度

Towards Full-parameter and Parameter-efficient Self-learning For Endoscopic Camera Depth Estimation

  • 分阶段优化注意力、卷积与MLP子空间,实现全参数适配
  • 在SCARED数据集上,误差指标下降至4.1%(原为10.2%)
  • 适合需要高精度深度估计的医疗视觉任务研究者

近期研究通过适配深度基础模型来解决内窥镜深度估计问题,但这类方法通常因限制参数搜索范围至低秩子空间而性能不足,并改变训练动态。为此,本文提出一种全参数且参数高效的自学习框架。第一阶段同时在不同子空间中适应注意力、卷积和多层感知机模块;第二阶段采用内存高效的优化策略,对子空间进行组合,进一步提升性能。在SCARED数据集上的初步实验表明,相较于现有最优模型,该方法在Sq Rel、Abs Rel、RMSE和RMSE log四项指标上分别从10.2%降至4.1%,显著提升精度。

原文摘要 · Abstract (English)

Adaptation methods are developed to adapt depth foundation models to endoscopic depth estimation recently. However, such approaches typically under-perform training since they limit the parameter search to a low-rank subspace and alter the training dynamics. Therefore, we propose a full-parameter and parameter-efficient learning framework for endoscopic depth estimation. At the first stage, the subspace of attention, convolution and multi-layer perception are adapted simultaneously within different sub-spaces. At the second stage, a memory-efficient optimization is proposed for subspace composition and the performance is further improved in the united sub-space. Initial experiments on the SCARED dataset demonstrate that results at the first stage improves the performance from 10.2% to 4.1% for Sq Rel, Abs Rel, RMSE and RMSE log in the comparison with the state-of-the-art models.

深度估计内窥镜自学习高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。