直接在观测空间生成3D物体,解决姿态对齐难题。
Pose-Aware Diffusion for 3D Generation

- 直接在观测空间生成3D几何,不依赖标准形态假设
- 相比顶尖方法,几何对齐与图像-3D对应更优
- 可自然组合生成场景,保持精确空间布局
由于解耦的标准化-旋转范式存在空间错位和变换歧义,生成姿态对齐的3D物体极具挑战。为此,我们提出Pose-Aware Diffusion(PAD),一种端到端扩散框架,直接在观测空间合成3D几何。通过将单目深度反投影为部分点云,并显式注入作为3D几何锚点,PAD摒弃了标准形态假设,实现严格的时空监督。这种原生生成机制从根本上解决了姿态歧义,生成高质量的姿态对齐资产。大量实验表明,PAD在几何对齐和图像-3D对应性上优于当前最先进方法。此外,通过独立生成物体的简单并集,PAD可自然扩展至组合式3D场景重建,凸显其保持精确空间布局的鲁棒能力。
原文摘要 · Abstract (English)
Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rotate paradigms. To this end, we introduce Pose-Aware Diffusion (PAD), a novel end-to-end diffusion framework that synthesizes 3D geometry directly within the observation space. By unprojecting monocular depth into a partial point cloud and explicitly injecting it as a 3D geometric anchor, PAD abandons canonical assumptions to enforce rigorous spatial supervision. This native generation intrinsically resolves pose ambiguity, producing high-fidelity pose-aligned assets. Extensive experiments demonstrate that PAD achieves superior geometric alignment and image-to-3D correspondence compared to state-of-the-art methods. Additionally, PAD naturally extends to compositional 3D scene reconstruction via a simple union of independently generated objects, highlighting its robust ability to preserve precise spatial layouts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。