用3D先验生成更真实连贯的轨道视频,解决背面视角难还原问题。
Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors
- 引入3D基础模型的全局与局部隐向量作为几何约束。
- 在多基准测试中视觉质量、形状真实性和多视角一致性均领先。
- 适合需要高保真3D结构生成的视频创作与虚拟展示场景。
我们提出一种新方法,仅凭单张图像生成几何上真实且一致的轨道视频。现有视频生成方法主要依赖像素级注意力维持视角一致性,但在长距离外推(如背面视角合成)时,因像素对应关系有限,难以保证结构合理。为此,我们利用3D基础生成模型中的丰富形状先验作为辅助约束,其通过大规模3D资产语料库学习到真实物体形状分布。具体地,我们使用3D基础模型编码的两个尺度的隐特征进行提示:(i) 去噪后的全局隐向量提供整体结构引导,(ii) 从体素特征投影出的一组隐图像提供视图依赖的细粒度几何细节。相比常用的2.5D表示(如深度图或法线图),这些紧凑特征可建模完整物体形状,并避免显式网格提取,提升推理效率。为实现有效形状条件化,我们设计多尺度3D适配器,通过交叉注意力将特征令牌注入基础视频模型,保留其通用视频预训练能力,并支持简单、模型无关的微调。大量实验表明,该方法在多个基准上优于当前最先进方法,在视觉质量、形状真实性和多视角一致性方面均有显著提升,且对复杂相机轨迹和真实图像具有强泛化能力。
原文摘要 · Abstract (English)
We present a novel method for generating geometrically realistic and consistent orbital videos from a single image of an object. Existing video generation works mostly rely on pixel-wise attention to enforce view consistency across frames. However, such mechanism does not impose sufficient constraints for long-range extrapolation, e.g. rear-view synthesis, in which pixel correspondences to the input image are limited. Consequently, these works often fail to produce results with a plausible and coherent structure. To tackle this issue, we propose to leverage rich shape priors from a 3D foundational generative model as an auxiliary constraint, motivated by its capability of modeling realistic object shape distributions learned from large 3D asset corpora. Specifically, we prompt the video generation with two scales of latent features encoded by the 3D foundation model: (i) a denoised global latent vector as an overall structural guidance, and (ii) a set of latent images projected from volumetric features to provide view-dependent and fine-grained geometry details. In contrast to commonly used 2.5D representations such as depth or normal maps, these compact features can model complete object shapes, and help to improve inference efficiency by avoiding explicit mesh extraction. To achieve effective shape conditioning, we introduce a multi-scale 3D adapter to inject feature tokens to the base video model via cross-attention, which retains its capabilities from general video pretraining and enables a simple and model-agonistic fine-tuning process. Extensive experiments on multiple benchmarks show that our method achieves superior visual quality, shape realism and multi-view consistency compared to state-of-the-art methods, and robustly generalizes to complex camera trajectories and in-the-wild images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。