仅用一张野外照片重建高精度可动画的狗3D模型
CORGI: Consistency-Aware 3D Dog Reconstruction from a Single Image in the Wild

- 通过规范姿态生成360度视频观测,解决单图歧义问题
- 利用神经变形场建模每视角生成误差,提升顶点级精度
- 自监督修复模块恢复高频细节,避免结构畸变
从一张无约束的野外图像中重建高保真、高度灵活的动物(如狗)3D模型仍是重大挑战。本文提出CORGI框架,完全无需3D监督即可实现一致性感知的3D狗重建。为克服生成不一致与缺乏多视角数据的问题,该方法引入三个核心组件:首先,提出基于规范姿态驱动的轨道生成(CDOG)策略,使用专用的规范与轨道LoRA对任意输入姿态进行归一化,并合成可靠的360度视频观测;其次,设计一致性感知可变形3DGS(CA-3DGS)模块,基于D-SMAL先验,通过专用神经变形场显式建模每视角生成误差,学习精确的顶点级位移;最后,引入自监督变形条件生成修复(DCGR)模块,消除结构畸变并恢复高频细节。大量实验表明,CORGI在多种犬种上均达到顶尖性能,生成几何准确、视觉连贯且完全可动画的3D资产,可直接用于下游应用。
原文摘要 · Abstract (English)
Reconstructing high-fidelity 3D models of highly articulated animals, such as dogs, from a single in-the-wild image remains a formidable challenge. In this paper, we introduce CORGI, a novel framework for consistency-aware 3D dog reconstruction from a single unconstrained image that completely eliminates the need for 3D supervision. To overcome generative inconsistencies and the lack of multi-view capture, our pipeline introduces three core components. First, we propose a Canonical-Driven Orbital Generation (CDOG) strategy, utilizing specialized Canonical and Orbit LoRAs to normalize arbitrary input poses and synthesize reliable 360-degree video observations. Second, we design a Consistency-aware Deformable 3DGS (CA-3DGS) module that anchors on a D-SMAL prior, explicitly modeling per-view generative errors through dedicated neural deformation fields to learn accurate vertex-level displacements. Finally, to eliminate structural distortions and recover high-frequency details, we introduce a self-supervised Deformation-Conditioned Generative Repair (DCGR) module. Extensive experiments demonstrate that CORGI achieves state-of-the-art performance, generalizing seamlessly across diverse dog breeds to produce geometrically accurate, visually coherent, and fully animatable 3D assets ready for downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。