arXiv:2601.05246cs.CV2026-01TPAMI被引 2

用像素空间生成模型,消除深度图中的漂浮像素,提升几何重建精度。

Pixel-Perfect Visual Geometry Estimation

  • 基于像素空间扩散变换器,通过语义提示与级联架构提升效率与细节。
  • 在单目与视频深度估计上均达当前最优,点云更干净无漂浮像素。
  • 适合机器人、增强现实等需高精度几何感知的场景使用。

从图像中恢复清晰准确的几何结构对机器人和增强现实至关重要。然而,现有几何基础模型仍严重受制于漂浮像素和细粒度细节丢失问题。本文提出像素级精确的视觉几何模型,通过在像素空间中利用生成建模实现高质量、无漂浮像素的点云预测。我们首先构建了基于像素空间扩散变换器(DiT)的单目深度基础模型Pixel-Perfect Depth(PPD)。为解决像素空间扩散带来的高计算复杂度,提出两项关键设计:1)语义提示扩散变换器(Semantics-Prompted DiT),引入视觉基础模型的语义表示以引导扩散过程,保留全局语义同时增强细粒度视觉细节;2)级联扩散变换器架构(Cascade DiT),逐步增加图像标记数,兼顾效率与精度。为进一步拓展至视频场景(PPVD),提出新的语义一致性扩散变换器,从多视角几何基础模型中提取时序一致的语义,并在DiT内执行参考引导的标记传播,以极低计算与内存开销维持时间连贯性。所提模型在所有生成式单目与视频深度估计模型中表现最佳,生成的点云显著优于其他模型。

原文摘要 · Abstract (English)

Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying pixels and the loss of fine details. In this paper, we present pixel-perfect visual geometry models that can predict high-quality, flying-pixel-free point clouds by leveraging generative modeling in the pixel space. We first introduce Pixel-Perfect Depth (PPD), a monocular depth foundation model built upon pixel-space diffusion transformers (DiT). To address the high computational complexity associated with pixel-space diffusion, we propose two key designs: 1) Semantics-Prompted DiT, which incorporates semantic representations from vision foundation models to prompt the diffusion process, preserving global semantics while enhancing fine-grained visual details; and 2) Cascade DiT architecture that progressively increases the number of image tokens, improving both efficiency and accuracy. To further extend PPD to video (PPVD), we introduce a new Semantics-Consistent DiT, which extracts temporally consistent semantics from a multi-view geometry foundation model. We then perform reference-guided token propagation within the DiT to maintain temporal coherence with minimal computational and memory overhead. Our models achieve the best performance among all generative monocular and video depth estimation models and produce significantly cleaner point clouds than all other models.

深度估计扩散模型几何重建视频生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。