用视频生成先验,把单图变3D高斯点云,还能防形变。
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
- 用分段小运动轨迹生成视频,解决大视角变化问题
- 通过MASt3R校准相机位姿,生成精确点云
- 设计畸变感知的高斯表示,保持3D结构一致性
单图像3D重建因几何模糊和视角信息有限仍具挑战。最近的潜在视频扩散模型(LVDMs)从大规模视频数据中学习了有前景的3D先验。然而有效利用这些先验面临三大难题:(1)大相机运动导致质量下降,(2)难以实现精确相机控制,(3)扩散过程固有的几何畸变破坏3D一致性。本文提出LiftImage3D框架,有效释放LVDM的生成先验并保障3D一致性。具体地,设计可操控的分段运动轨迹生成视频帧,将大运动分解为可控小运动;使用鲁棒神经匹配模型MASt3R校准生成帧的相机位姿,构建对应点云;最后提出畸变感知的3D高斯点阵表示,可独立学习帧间畸变并输出无畸变的标准高斯点。大量实验表明,LiftImage3D在LLFF、DL3DV和Tanks and Temples三个挑战性数据集上达到领先性能,并能良好泛化至多样真实图像,涵盖卡通插画到复杂真实场景。
原文摘要 · Abstract (English)
Single-image 3D reconstruction remains a fundamental challenge in computer vision due to inherent geometric ambiguities and limited viewpoint information. Recent advances in Latent Video Diffusion Models (LVDMs) offer promising 3D priors learned from large-scale video data. However, leveraging these priors effectively faces three key challenges: (1) degradation in quality across large camera motions, (2) difficulties in achieving precise camera control, and (3) geometric distortions inherent to the diffusion process that damage 3D consistency. We address these challenges by proposing LiftImage3D, a framework that effectively releases LVDMs' generative priors while ensuring 3D consistency. Specifically, we design an articulated trajectory strategy to generate video frames, which decomposes video sequences with large camera motions into ones with controllable small motions. Then we use robust neural matching models, i.e. MASt3R, to calibrate the camera poses of generated frames and produce corresponding point clouds. Finally, we propose a distortion-aware 3D Gaussian splatting representation, which can learn independent distortions between frames and output undistorted canonical Gaussians. Extensive experiments demonstrate that LiftImage3D achieves state-of-the-art performance on two challenging datasets, i.e. LLFF, DL3DV, and Tanks and Temples, and generalizes well to diverse in-the-wild images, from cartoon illustrations to complex real-world scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。