arXiv:2412.14015cs.CV2024-12CVPR被引 102

用低成本激光雷达作提示,实现4K级高精度深度图生成

Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation

论文配图:Prompting Depth Anything for 4K Resolution Accurate Metric Depth Estimation
图 1 · 摘自论文原文
  • 以激光雷达数据作为提示,融合多尺度信息指导深度模型推理
  • 在ARKitScenes和ScanNet++上达到新基准,支持4K分辨率输出
  • 适合需要高精度深度的3D重建与机器人抓取等应用

提示在激发视觉语言基础模型潜力中起关键作用。本文首次将提示引入深度基础模型,提出名为Prompt Depth Anything的新范式,用于精准度量深度估计。通过使用低成本激光雷达作为提示,引导Depth Anything模型输出高精度度量深度,最高可达4K分辨率。方法核心在于一种简洁的提示融合设计,在深度解码器中多尺度融合激光雷达信息。针对同时包含激光雷达深度与精确真值深度的训练数据稀缺问题,提出可扩展的数据管道,包含合成数据的激光雷达模拟和真实数据伪真值深度生成。该方法在ARKitScenes和ScanNet++数据集上取得新最优性能,并推动3D重建与泛化机器人抓取等下游应用的发展。

原文摘要 · Abstract (English)

Prompts play a critical role in unleashing the power of language and vision foundation models for specific tasks. For the first time, we introduce prompting into depth foundation models, creating a new paradigm for metric depth estimation termed Prompt Depth Anything. Specifically, we use a low-cost LiDAR as the prompt to guide the Depth Anything model for accurate metric depth output, achieving up to 4K resolution. Our approach centers on a concise prompt fusion design that integrates the LiDAR at multiple scales within the depth decoder. To address training challenges posed by limited datasets containin both LiDAR depth and precise GT depth, we propose a scalable data pipeline that includes synthetic data LiDAR simulation and real data pseudo GT depth generation. Our approach sets new state-of-the-arts on the ARKitScenes and ScanNet++ datasets and benefits downstream applications, including 3D reconstruction and generalized robotic grasping.

深度估计提示学习4K分辨率激光雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。