用2D先验+3D约束,生成遮挡下完整且真实的3D模型。
GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
- 融合2D生成先验与3D几何推理,实现多视角一致的3D重建。
- 在合成与真实场景中均优于现有方法,保持几何准确性。
- 适合需要高保真3D生成的工业设计、自动驾驶等场景。
在部分遮挡(即非模态)情况下生成完整的3D物体是实际应用中重要但具挑战性的问题,因为现实场景中大量几何信息不可见。现有方法要么直接在3D空间操作,虽保证几何一致性但生成表达力弱;要么依赖2D非模态补全,虽有强外观先验但难以保证可靠3D结构。为此,本文提出GENA3D框架,通过条件3D生成范式整合学习的2D生成先验与显式的3D几何推理。2D先验使模型能合理推断多样化的遮挡内容,3D表示则确保多视图一致性与空间有效性。设计引入新颖的视图级交叉注意力实现多视角对齐,以及立体条件交叉注意力将生成预测锚定于3D关系中。结合生成想象与结构约束,GENA3D可从有限观测生成完整且连贯的3D物体,不牺牲几何保真度。实验表明,该方法在合成与真实世界非模态场景中均优于现有方法,验证了融合2D先验与3D一致性在复杂环境中生成可信且几何一致3D结构的有效性。
原文摘要 · Abstract (English)
Generating complete 3D objects under partial occlusions (i.e., amodal scenarios) is a practically important yet challenging problem, as large portions of object geometry are unobserved in real-world scenarios. Existing approaches either operate directly in 3D, which ensures geometric consistency but often lacks generative expressiveness, or rely on 2D amodal completion, which provides strong appearance priors but does not guarantee reliable 3D structure. This raises a key question: how can we achieve both generative plausibility and geometric coherence in amodal 3D modeling? To answer this question, we introduce GENA3D (GENarative Amodal 3D), a framework that integrates learned 2D generative priors with explicit 3D geometric reasoning within a conditional 3D generation paradigm. The 2D priors enable the model to plausibly infer diverse occluded content, while the 3D representation enforces multi-view consistency and spatial validity. Our design incorporates a novel View-Wise Cross-Attention for multi-view alignment and a Stereo-Conditioned Cross-Attention to anchor generative predictions in 3D relationships. By combining generative imagination with structural constraints, GENA3D generates complete and coherent 3D objects from limited observations without sacrificing geometric fidelity. Experiments demonstrate that our method outperforms existing approaches in both synthetic and real-world amodal scenarios, highlighting the effectiveness of bridging 2D priors and 3D coherence in generating plausible and geometrically consistent 3D structures in complex environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。