arXiv:2607.08771cs.CV2026-07

轻量级单目深度模型ZipDepth,实现跨域实时推理。

ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device

论文配图:ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device
图 1 · 摘自论文原文
  • 用可重参数化编码器-解码器+大模型知识蒸馏
  • 610万参数,5个基准上轻量模型最佳精度-效率平衡
  • 适合移动端和嵌入式设备部署,跨域表现强

单目深度估计虽因基础模型取得显著进展,具备出色的零样本泛化能力,但其计算开销远超嵌入式与移动平台的承受范围。现有轻量级方案多限于单一领域、自监督范式,在领域迁移下表现不佳。本文提出ZipDepth,一种紧凑的单目深度网络,通过结合高效可重参数化编码器-解码器结构,并在大规模多领域数据集上从基础模型进行知识蒸馏,实现性能突破。该模型仅含610万参数,可在从服务器GPU到低功耗设备的多种平台上实现实时运行,在五个基准测试中达到轻量模型最优的零样本准确率与部署效率平衡,以50倍更少参数逼近基础模型精度。

原文摘要 · Abstract (English)

Monocular depth estimation has seen remarkable progress through foundation models achieving robust zero-shot generalization, yet their computational demands place them far beyond the reach of embedded and mobile platforms. Lightweight alternatives exist, but have been developed almost exclusively within single-domain, self-supervised paradigms, failing silently under domain shift. We present ZipDepth, a compact monocular depth network that bridges this gap by combining an efficient reparameterizable encoder-decoder with large-scale knowledge distillation from a foundation model over a large multi-domain training set. Comprising just 6.1M parameters, ZipDepth runs at real-time rates from server GPUs to power-constrained devices, achieving the best trade-off between zero-shot accuracy and deployment efficiency among lightweight models across five benchmarks, taking a significant step towards the accuracy of foundation models with 50x more parameters.

单目深度轻量化知识蒸馏边缘部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。