arXiv:2603.21557cs.CV2026-03中稿 · ICME 2026被引 1

自适应层次结构让3D生成模型能自动发现物体部件,通用性强且不依赖人工设定部件数。

From Part to Whole: 3D Generative World Model with an Adaptive Structural Hierarchy

  • 通过动态激活机制在隐空间自动发现部件,无需预设部件数量。
  • 在跨类别迁移和新布局泛化上显著优于现有方法,部件数可灵活扩展。
  • 适合需要通用3D生成、避免人工标注部件的科研与工业场景。

单图像3D生成是视觉到图形模型的核心任务,但在稀疏监督下,如何在多样语义类别和复杂结构间实现可靠泛化仍具挑战。现有方法多采用固定结构或预设部件数,如PartCrafter需人工指定部件数量,易导致过拟合、结构碎片化或缺失,且难以泛化至新布局。本文提出一种从局部到整体的自适应层次结构3D生成世界模型,通过图像标记直接推断软性、组合式掩码,自主发现隐空间中的结构槽。设计自适应槽门机制,动态调节各槽激活概率,并平滑合并冗余槽,确保结构紧凑而表达丰富。每个提取槽对齐可学习的类别无关原型库,利用通用几何原型实现跨类别形状共享与去噪。引入轻量级3D去噪器,通过统一扩散目标重建几何与外观。实验显示,在跨类别迁移与部件数外推上持续提升,消融实验证明原型库支持形状先验共享,槽门机制促进结构自适应。

原文摘要 · Abstract (English)

Single-image 3D generation lies at the core of vision-to-graphics models in the real world. However, it remains a fundamental challenge to achieve reliable generalization across diverse semantic categories and highly variable structural complexity under sparse supervision. Existing approaches typically model objects in a monolithic manner or rely on a fixed number of parts, including recent part-aware models such as PartCrafter, which still require a labor-intensive user-specified part count. Such designs easily lead to overfitting, fragmented or missing structural components, and limited compositional generalization when encountering novel object layouts. To this end, this paper rethinks single-image 3D generation as learning an adaptive part-whole hierarchy in the flexible 3D latent space. We present a novel part-to-whole 3D generative world model that autonomously discovers latent structural slots by inferring soft and compositional masks directly from image tokens. Specifically, an adaptive slot-gating mechanism dynamically determines the slot-wise activation probabilities and smoothly consolidates redundant slots within different objects, ensuring that the emergent structure remains compact yet expressive across categories. Each distilled slot is then aligned to a learnable, class-agnostic prototype bank, enabling powerful cross-category shape sharing and denoising through universal geometric prototypes in the real world. Furthermore, a lightweight 3D denoiser is introduced to reconstruct geometry and appearance via unified diffusion objectives. Experiments show consistent gains in cross-category transfer and part-count extrapolation, and ablations confirm complementary benefits of the prototype bank for shape-prior sharing as well as slot-gating for structural adaptation.

3D生成自适应结构扩散模型原型共享

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。