Cupid通过联合建模物体与相机位姿,实现高保真3D生成重建。
CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- 分两阶段生成:先粗略结构与位姿估计,再注入图像特征精细优化
- 在PSNR上超现有方法3 dB,Chamfer Distance提升10%
- 可直接扩展至多视角和场景级重建,无需后期优化
我们提出Cupid,一种联合建模规范物体与相机位姿完整分布的生成式3D重建框架。其基于流模型的两阶段流程首先生成粗略3D结构与2D-3D对应关系,以鲁棒估计相机位姿;随后,在该位姿条件下,将像素对齐的图像特征直接注入生成过程,融合生成模型的先验与重建的几何精度。该策略实现了卓越的保真度,在PSNR上超越当前最优方法超过3 dB,Chamfer Distance提升10%。作为解耦物体与相机位姿的统一生成模型,Cupid可自然扩展至多视角及场景级重建任务,无需后处理优化或微调。
原文摘要 · Abstract (English)
We introduce Cupid, a generative 3D reconstruction framework that jointly models the full distribution over both canonical objects and camera poses. Our two-stage flow-based model first generates a coarse 3D structure and 2D-3D correspondences to estimate the camera pose robustly. Conditioned on this pose, a refinement stage injects pixel-aligned image features directly into the generative process, marrying the rich prior of a generative model with the geometric fidelity of reconstruction. This strategy achieves exceptional faithfulness, outperforming state-of-the-art reconstruction methods by over 3 dB PSNR and 10% in Chamfer Distance. As a unified generative model that decouples the object and camera pose, Cupid naturally extends to multi-view and scene-level reconstruction tasks without requiring post-hoc optimization or fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。