arXiv:2603.27533cs.CVcs.AI2026-03中稿 · ICASSP 2026, 5 pag…

融合深度与单目语义信息,提升无CAD模型下的物体6D姿态估计精度。

Demo-Pose: Depth-Monocular Modality Fusion For Object Pose Estimation

  • 用新融合策略结合单目语义与深度图的图卷积特征
  • 在REAL275上3D IoU提升3.2%,姿态准确率提高11.1%
  • 适合机器人抓取、AR/VR等需实时3D姿态的应用

物体姿态估计是3D视觉的基础任务,广泛应用于机器人、AR/VR和场景理解。本文解决从RGB-D输入中进行类别级9自由度姿态估计(6D姿态+3D尺寸)的问题,推理时不依赖CAD模型。现有深度仅方法虽表现强,但忽略RGB语义信息;多数RGB-D融合模型因跨模态融合不佳,未能对齐语义与几何表征而性能受限。我们提出DeMo-Pose,一种混合架构,通过新颖的多模态融合策略,将单目语义特征与基于深度的图卷积表示相融合。为进一步增强几何推理,引入无需推理开销的网格点损失(Mesh-Point Loss, MPL)。本方法实现实时推理,在多个物体类别上显著优于现有最优方法,在REAL275基准上相较强基线GPV-Pose提升3.2%的3D IoU和11.1%的姿态准确率。结果表明深度-RGB融合与几何感知学习的有效性,为真实世界应用提供鲁棒的类别级3D姿态估计能力。

原文摘要 · Abstract (English)

Object pose estimation is a fundamental task in 3D vision with applications in robotics, AR/VR, and scene understanding. We address the challenge of category-level 9-DoF pose estimation (6D pose + 3Dsize) from RGB-D input, without relying on CAD models during inference. Existing depth-only methods achieve strong results but ignore semantic cues from RGB, while many RGB-D fusion models underperform due to suboptimal cross-modal fusion that fails to align semantic RGB cues with 3D geometric representations. We propose DeMo-Pose, a hybrid architecture that fuses monocular semantic features with depth-based graph convolutional representations via a novel multimodal fusion strategy. To further improve geometric reasoning, we introduce a novel Mesh-Point Loss (MPL) that leverages mesh structure during training without adding inference overhead. Our approach achieves real-time inference and significantly improves over state-of-the-art methods across object categories, outperforming the strong GPV-Pose baseline by 3.2\% on 3D IoU and 11.1\% on pose accuracy on the REAL275 benchmark. The results highlight the effectiveness of depth-RGB fusion and geometry-aware learning, enabling robust category-level 3D pose estimation for real-world applications.

姿态估计多模态融合实时推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。