arXiv:2608.28549cs.CVcs.AI2026-08

用视频生成模型学几何,少数据也能超好效果。

Video Generative Models as Geometry Learner

论文配图:Video Generative Models as Geometry Learner
图 1 · 摘自论文原文
  • 把视频生成模型改造成几何预测任务,以未来帧预测为思路
  • 零样本下在多数据集上超越现有生成与判别方法,用数据少100倍
  • 适合资源有限但追求高精度几何估计的研究者

近期的几何估计生成方法多基于预训练图像扩散模型,将任务视为图像条件生成。现有方法要么独立训练特定几何模型(如深度、法向估计),忽略几何目标间的内在关联;要么联合微调修改后的图像扩散主干(如调整自注意力机制),通常需要大量标注数据。为克服这些局限,我们创新性地将预训练视频生成模型重新用于统一且数据高效的几何估计框架,将其建模为未来帧预测任务。所提方法GeoNeXt继承视频模型固有的结构化知识与更丰富的先验,进一步适配图像与几何目标之间的联合建模(图像 <-> 几何),实现更高效的数据利用与几何学习。大量实验验证了该方法在跨多个数据集的零样本单目深度与表面法向估计中表现优异,优于此前任务特定及统一生成方法,且训练数据显著减少。值得注意的是,本方法性能媲美需超100倍数据训练的判别式最先进方法,甚至在部分基准上更优。

原文摘要 · Abstract (English)

Recent generative approaches to geometry estimation adapt pretrained image diffusion models and treat the task as image-conditioned generation. Leveraging off-the-shelf image diffusion models, they either (i) train task-specific geometry models (for depth and surface normal estimation) independently, losing the opportunity of exploring the intrinsic correlation of these geometric targets, or (ii) jointly fine-tune modified image diffusion backbones (e.g., altered self-attention), which typically demands substantial labeled data. To overcome these limitations in a principled fashion, we repurpose pretrained video generative models as a unified and data-efficient framework for geometry estimation, formulated innovatively as a next-frames prediction task. Our method, GeoNeXt, inherits naturally structured knowledge and richer priors from the video model, while further adapting them for joint modeling of images and geometry targets (image <-> geometry), enabling more data efficient and effective learning of geometry. Extensive experiments validate our method for zero-shot monocular depth and surface normal estimation across diverse datasets, outperforming both previous task-specific and unified generative competitors while using substantially less training data. Notably, our method rivals discriminative state-of-the-art approaches trained on over 100x more data and even standouts on several benchmarks.

几何估计视频生成扩散模型少样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。