用视频扩散模型统一估计全局几何,实现跨帧一致性。
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
- 以全局坐标系为基准,预测与视频帧对齐的几何属性。
- 重用位置编码设计高效条件输入,提升跨帧一致性。
- 多属性联合训练,静态数据也能泛化到动态场景。
近期,利用扩散模型先验辅助单目几何估计(如深度、法向)的方法因具备强泛化能力而备受关注。然而,多数现有方法仅关注单帧内相机坐标系下的几何属性估计,忽略了扩散模型在帧间对应关系上的内在能力。本文通过合理设计与微调,有效利用视频生成模型的固有连续性,实现一致的几何估计。具体包括:1)选择与视频帧具有相同对应关系的全局坐标系几何属性作为预测目标;2)通过重用位置编码提出一种新颖高效的条件控制方法;3)通过共享对应关系的多个几何属性联合训练提升性能。实验表明,该方法在视频中预测全局几何属性上表现优越,可直接应用于重建任务。即使仅在静态视频数据上训练,仍具备向动态视频场景泛化的能力。
原文摘要 · Abstract (English)
Recently, methods leveraging diffusion model priors to assist monocular geometric estimation (e.g., depth and normal) have gained significant attention due to their strong generalization ability. However, most existing works focus on estimating geometric properties within the camera coordinate system of individual video frames, neglecting the inherent ability of diffusion models to determine inter-frame correspondence. In this work, we demonstrate that, through appropriate design and fine-tuning, the intrinsic consistency of video generation models can be effectively harnessed for consistent geometric estimation. Specifically, we 1) select geometric attributes in the global coordinate system that share the same correspondence with video frames as the prediction targets, 2) introduce a novel and efficient conditioning method by reusing positional encodings, and 3) enhance performance through joint training on multiple geometric attributes that share the same correspondence. Our results achieve superior performance in predicting global geometric attributes in videos and can be directly applied to reconstruction tasks. Even when trained solely on static video data, our approach exhibits the potential to generalize to dynamic video scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。