arXiv:2512.01030cs.CV2025-12被引 17

用扩散模型先验实现高精度单图几何预测,仅需少量数据即达顶尖水平。

Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model

  • 两阶段设计:先生成全局结构,再通过确定性流匹配细化细节。
  • 仅用5.9万样本训练,在深度估计上超越现有方法。
  • 适合需要少样本、高精度几何推理的研究与应用。

从单张图像恢复像素级几何属性本质上是病态问题,源于外观模糊和2D观测与3D结构之间的非单射映射。尽管判别式回归模型在大规模监督下表现强劲,其性能受限于数据规模、质量和多样性,以及物理推理能力不足。近期扩散模型展现出强大的世界先验,编码了从海量图文数据中学到的几何与语义信息,但直接复用其随机生成范式对确定性几何推断并不理想:前者优化目标为多样化且高保真图像生成,而后者需要稳定准确的预测。本文提出Lotus-2,一种两阶段确定性框架,旨在稳定、准确、细粒度地完成几何密集预测,并提供最优适配协议以充分挖掘预训练生成先验。第一阶段中,核心预测器采用单步确定性形式,结合干净数据目标与轻量局部连续性模块(LCM),生成无网格伪影的全局一致结构。第二阶段,细节锐化器在核心预测器定义的流形内执行约束多步修正流重构,通过无噪声确定性流匹配增强细粒度几何。仅使用5.9万训练样本(不足现有大数据集的1%),Lotus-2在单目深度估计上建立新SOTA,表面法向预测达到极具竞争力的水平。结果表明,扩散模型可作为确定性世界先验,实现超越传统判别与生成范式的高质量几何推理。

原文摘要 · Abstract (English)

Recovering pixel-wise geometric properties from a single image is fundamentally ill-posed due to appearance ambiguity and non-injective mappings between 2D observations and 3D structures. While discriminative regression models achieve strong performance through large-scale supervision, their success is bounded by the scale, quality, and diversity of available data, as well as by limited physical reasoning. Recent diffusion models exhibit powerful world priors that encode geometry and semantics learned from massive image-text data, yet directly reusing their stochastic generative formulation is suboptimal for deterministic geometric inference: the former is optimized for diverse and high-fidelity image generation, whereas the latter requires stable and accurate predictions. In this work, we propose Lotus-2, a two-stage deterministic framework for stable, accurate and fine-grained geometric dense prediction, aiming to provide an optimal adaptation protocol to fully exploit the pre-trained generative priors. Specifically, in the first stage, the core predictor employs a single-step deterministic formulation with a clean-data objective and a lightweight local continuity module (LCM) to generate globally coherent structures without grid artifacts. In the second stage, the detail sharpener performs a constrained multi-step rectified-flow refinement within the manifold defined by the core predictor, enhancing fine-grained geometry through noise-free deterministic flow matching. Using only 59K training samples, less than 1% of existing large-scale datasets, Lotus-2 establishes new state-of-the-art results in monocular depth estimation and highly competitive surface normal prediction. These results demonstrate that diffusion models can serve as deterministic world priors, enabling high-quality geometric reasoning beyond traditional discriminative and generative paradigms.

几何推理扩散模型单目深度少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。