arXiv:2511.14238cs.CVcs.LG2025-11中稿 · AAAI被引 1

用弱监督自训练提升单目深度估计模型泛化能力

Enhancing Generalization of Depth Estimation Foundation Model via Weakly-Supervised Adaptation with Regularization

  • 通过密集自训练与层级语义归一化增强结构稳定性
  • 引入成对序数深度标注,减少局部拓扑错误
  • 轻量级正则化保持模型通用知识,适合部署到新场景

基础模型的出现显著提升了单目深度估计(MDE)的零样本泛化能力,如Depth Anything系列所示。然而,当有下游任务数据可用时,能否进一步提升性能?为此,我们提出WeSTAR——一种参数高效、基于弱监督自训练与正则化的适应框架,旨在增强MDE基础模型在未见多样域中的鲁棒性。首先采用密集自训练目标作为主要结构自监督信号;为提升鲁棒性,引入语义感知的分层归一化,利用实例级分割图实现更稳定、多尺度的结构归一化。此外,通过成对序数深度标注提供低成本弱监督,施加信息性序数约束以缓解局部拓扑误差。最后,采用权重正则化损失锚定LoRA更新,确保训练稳定并保留模型可泛化知识。在真实与损坏的分布外数据集上,多种挑战性场景下的大量实验表明,WeSTAR持续提升泛化性能,在广泛基准上达到最先进水平。

原文摘要 · Abstract (English)

The emergence of foundation models has substantially advanced zero-shot generalization in monocular depth estimation (MDE), as exemplified by the Depth Anything series. However, given access to some data from downstream tasks, a natural question arises: can the performance of these models be further improved? To this end, we propose WeSTAR, a parameter-efficient framework that performs Weakly supervised Self-Training Adaptation with Regularization, designed to enhance the robustness of MDE foundation models in unseen and diverse domains. We first adopt a dense self-training objective as the primary source of structural self-supervision. To further improve robustness, we introduce semantically-aware hierarchical normalization, which exploits instance-level segmentation maps to perform more stable and multi-scale structural normalization. Beyond dense supervision, we introduce a cost-efficient weak supervision in the form of pairwise ordinal depth annotations to further guide the adaptation process, which enforces informative ordinal constraints to mitigate local topological errors. Finally, a weight regularization loss is employed to anchor the LoRA updates, ensuring training stability and preserving the model's generalizable knowledge. Extensive experiments on both realistic and corrupted out-of-distribution datasets under diverse and challenging scenarios demonstrate that WeSTAR consistently improves generalization and achieves state-of-the-art performance across a wide range of benchmarks.

深度估计自训练弱监督泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。