从单图生成可交互的3D组合物体,解决遮挡下的几何失真问题。
Interact3D: Compositional 3D Generation of Interactive Objects
- 基于统一3D引导场景,分阶段融合物体并优化空间关系。
- 通过SDF优化与自纠正机制,减少碰撞并提升几何保真度。
- 适合需要高精度3D场景生成的研究者和工业设计人员。
近期3D生成技术已实现高质量单体资产合成,但仅凭单张图像生成具有相互作用的3D组合物体——尤其在存在遮挡时——仍具挑战性。现有方法常导致隐藏区域几何细节退化,且无法保持物体间的空间关系(OOR)。本文提出新框架Interact3D,旨在生成物理合理的交互式3D组合物体。首先利用先进生成先验,在统一3D引导场景中构建高质量单体资产。随后引入稳健的两阶段组合流程:主物体通过精确的全局-局部几何对齐(注册)锚定,其余物体则通过基于可微SDF的优化整合,显式惩罚几何相交。为缓解碰撞难题,进一步部署闭环、代理式精炼策略:视觉-语言模型(VLM)自动分析多视角渲染结果,生成针对性修正提示,并引导图像编辑模块迭代自校正生成流程。大量实验表明,Interact3D成功生成具备碰撞感知能力的组合,显著提升几何保真度与空间关系一致性。
原文摘要 · Abstract (English)
Recent breakthroughs in 3D generation have enabled the synthesis of high-fidelity individual assets. However, generating 3D compositional objects from single images--particularly under occlusions--remains challenging. Existing methods often degrade geometric details in hidden regions and fail to preserve the underlying object-object spatial relationships (OOR). We present a novel framework Interact3D designed to generate physically plausible interacting 3D compositional objects. Our approach first leverages advanced generative priors to curate high-quality individual assets with a unified 3D guidance scene. To physically compose these assets, we then introduce a robust two-stage composition pipeline. Based on the 3D guidance scene, the primary object is anchored through precise global-to-local geometric alignment (registration), while subsequent geometries are integrated using a differentiable Signed Distance Field (SDF)-based optimization that explicitly penalizes geometry intersections. To reduce challenging collisions, we further deploy a closed-loop, agentic refinement strategy. A Vision-Language Model (VLM) autonomously analyzes multi-view renderings of the composed scene, formulates targeted corrective prompts, and guides an image editing module to iteratively self-correct the generation pipeline. Extensive experiments demonstrate that Interact3D successfully produces promising collsion-aware compositions with improved geometric fidelity and consistent spatial relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。