用多模态信号统一完成3D形状补全,支持遮挡和编辑场景。
Axolotl3D: a Unified Framework for Faithful 3D Shape Completion

- 融合图像、可见性掩码、相机参数与部分点云进行联合建模
- 在Toys4K和OmniObject3D上实现最佳补全精度,遮挡下仍保持稳定
- 适合需要真实重建与可控编辑的3D生成任务
近期3D生成模型利用大规模先验和扩散架构,从单张图像生成高质量几何体。然而,它们假设视角完整且仅输入单视图,难以应用于多视角、遮挡或编辑场景。尽管已有研究分别解决这些问题,但缺乏统一框架来应对多样条件下的可控3D补全。我们提出Axolotl3D,一种多模态且考虑遮挡的3D生成模型,同时以图像、可见性掩码、相机参数和部分点云作为条件。点云作为几何锚点,促进形状补全的忠实性;相机参数确保在共享3D坐标系中多视角一致性。统一训练策略从大规模3D数据合成多种条件场景,实现鲁棒的跨模态推理。在Toys4K和OmniObject3D上的实验表明,该方法在干净和遮挡环境下均达到最先进性能,并在真实世界重建与几何一致编辑中表现优异。
原文摘要 · Abstract (English)
Recent 3D generative models produce high-quality geometry from a single image using large-scale priors and diffusion architectures. However, they assume complete visibility and single-view inputs, limiting applicability in multi-view, occluded, or editing scenarios. Although prior works address these challenges individually, they lack a unified framework for controllable 3D completion under diverse conditioning signals. We present Axolotl3D, a multi-modal and occlusion-aware 3D generation model that jointly conditions on images, visibility masks, camera parameters, and a partial point cloud. The point cloud serves as a geometric anchor promoting faithful shape completion, while camera parameters ensure consistent multi-view alignment in a shared 3D coordinate system. A unified training strategy synthesizes diverse conditioning regimes from large-scale 3D data, enabling robust cross-modal reasoning. Experiments on Toys4K and OmniObject3D demonstrate state-of-the-art performance under both clean and occluded settings, as well as strong results in real-world reconstruction and geometry-consistent editing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。