arXiv:2605.00345cs.CV2026-05

直接在观测空间生成3D物体,解决姿态对齐难题。

Pose-Aware Diffusion for 3D Generation

论文配图:Pose-Aware Diffusion for 3D Generation
图 1 · 摘自论文原文
  • 直接在观测空间生成3D几何,不依赖标准形态假设
  • 相比顶尖方法,几何对齐与图像-3D对应更优
  • 可自然组合生成场景,保持精确空间布局

由于解耦的标准化-旋转范式存在空间错位和变换歧义,生成姿态对齐的3D物体极具挑战。为此,我们提出Pose-Aware Diffusion(PAD),一种端到端扩散框架,直接在观测空间合成3D几何。通过将单目深度反投影为部分点云,并显式注入作为3D几何锚点,PAD摒弃了标准形态假设,实现严格的时空监督。这种原生生成机制从根本上解决了姿态歧义,生成高质量的姿态对齐资产。大量实验表明,PAD在几何对齐和图像-3D对应性上优于当前最先进方法。此外,通过独立生成物体的简单并集,PAD可自然扩展至组合式3D场景重建,凸显其保持精确空间布局的鲁棒能力。

原文摘要 · Abstract (English)

Generating pose-aligned 3D objects is challenging due to the spatial mismatches and transformation ambiguities inherent in decoupled canonical-then-rotate paradigms. To this end, we introduce Pose-Aware Diffusion (PAD), a novel end-to-end diffusion framework that synthesizes 3D geometry directly within the observation space. By unprojecting monocular depth into a partial point cloud and explicitly injecting it as a 3D geometric anchor, PAD abandons canonical assumptions to enforce rigorous spatial supervision. This native generation intrinsically resolves pose ambiguity, producing high-fidelity pose-aligned assets. Extensive experiments demonstrate that PAD achieves superior geometric alignment and image-to-3D correspondence compared to state-of-the-art methods. Additionally, PAD naturally extends to compositional 3D scene reconstruction via a simple union of independently generated objects, highlighting its robust ability to preserve precise spatial layouts.

3D生成扩散模型姿态对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。