arXiv:2603.01371cs.CV2026-03

无需训练即可生成高空间保真度的多物体3D模型

TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity

  • 通过早期去噪阶段的实例感知分离引导,实现物体解耦
  • 引入几何自适应更新模块,保持物体相对位置与形状特征
  • 无需训练、推理更快,适合真实场景的3D内容生成

图像到3D多实例生成中的精确空间保真度对下游实际应用至关重要。现有方法通常在多实例数据集上微调预训练的图像到3D(I23D)模型,但带来巨大训练开销且难以保证空间一致性。我们观察到,预训练的I23D模型已具备有意义的空间先验,但因实例纠缠问题未被充分利用。为此,我们提出TIMI——一种无需训练的图像到3D多实例生成新框架,可实现高空间保真度。具体而言,我们引入实例感知分离引导(ISG)模块,在早期去噪阶段促进实例解耦;为进一步稳定ISG引导,设计空间稳定几何自适应更新(SGU)模块,以保留实例几何特征并维持其相对关系。大量实验表明,相比现有方法,本方法在全局布局和局部实例区分度上表现更优,且无需额外训练,推理速度更快。

原文摘要 · Abstract (English)

Precise spatial fidelity in Image-to-3D multi-instance generation is critical for downstream real-world applications. Recent work attempts to address this by fine-tuning pre-trained Image-to-3D (I23D) models on multi-instance datasets, which incurs substantial training overhead and struggles to guarantee spatial fidelity. In fact, we observe that pre-trained I23D models already possess meaningful spatial priors, which remain underutilized as evidenced by instance entanglement issues. Motivated by this, we propose TIMI, a novel Training-free framework for Image-to-3D Multi-Instance generation that achieves high spatial fidelity. Specifically, we first introduce an Instance-aware Separation Guidance (ISG) module, which facilitates instance disentanglement during the early denoising stage. Next, to stabilize the guidance introduced by ISG, we devise a Spatial-stabilized Geometry-adaptive Update (SGU) module that promotes the preservation of the geometric characteristics of instances while maintaining their relative relationships. Extensive experiments demonstrate that our method yields better performance in terms of both global layout and distinct local instances compared to existing multi-instance methods, without requiring additional training and with faster inference speed.

3D生成无训练空间保真实例解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。