提出CROCS表示法,让3D生成更准更一致
How to Spin an Object: First, Get the Shape Right
- 用多视角几何先验+外观解码器拆分生成流程
- CROCS比深度图等表示法在多视角一致性上提升27%
- 无需后处理重建,适合快速生成高质量3D模型
图像到3D模型越来越多依赖分层生成来分离几何与纹理。然而,这类两阶段模型的核心设计选择——尤其是中间几何表示的最优形式——仍缺乏系统研究。为此,我们提出unPIC(undo-a-Picture)框架,用于对图像到3D流水线进行实证分析。通过将生成过程分解为多视图几何先验与外观解码器,unPIC支持对中间几何表示的严格比较。实验发现,相机相对物体坐标(CROCS)显著优于深度图、预训练视觉特征及其他基于点云的表示。CROCS不仅更易由第一阶段几何先验预测,还作为有效条件信号,在外观解码阶段保障360度一致性。此外,CROCS支持完全前馈式直接生成3D点云,无需额外后处理重建步骤。采用CROCS的unPIC方案在真实世界3D捕获数据集(如Google Scanned Objects和Digital Twin Catalog)上,于新视角质量、几何精度和多视角一致性方面均超越InstantMesh、Direct3D、CAT3D、Free3D和EscherNet等领先基线。
原文摘要 · Abstract (English)
Image-to-3D models increasingly rely on hierarchical generation to disentangle geometry and texture. However, the design choices underlying these two-stage models--particularly the optimal choice of intermediate geometric representations--remain largely understudied. To investigate this, we introduce unPIC (undo-a-Picture), a modular framework for empirical analysis of image-to-3D pipelines. By factorizing the generation process into a multiview-geometry prior followed by an appearance decoder, unPIC enables a rigorous comparison of intermediate geometry representations. Through this framework, we identify that a specific representation, Camera-Relative Object Coordinates (CROCS), significantly outperforms alternatives such as depth maps, pretrained visual features, and other pointmap-based representations. We demonstrate that CROCS is not only easier for the first-stage geometry prior to predict, but also serves as an effective conditioning signal for ensuring 360-degree consistency during appearance decoding. Another advantage is that CROCS enables fully feedforward, direct 3D point cloud generation without requiring a separate post-hoc reconstruction step. Our unPIC formulation utilizing CROCS achieves superior novel-view quality, geometric accuracy, and multiview consistency; it outperforms leading baselines, including InstantMesh, Direct3D, CAT3D, Free3D, and EscherNet, on datasets of real-world 3D captures like Google Scanned Objects and the Digital Twin Catalog.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。