OMNI-PoseX实现快速6D姿态估计,支持开放世界物体识别与稳定抓取。
OMNI-PoseX: A Fast Vision Model for 6D Object Pose Estimation in Embodied Tasks
- 将语义理解与旋转推理解耦,用轻量多模态融合提升精度
- 在多个基准上达到最新最优性能,推理速度满足实时需求
- 适合机器人抓取、开放世界感知等需要泛化能力的场景
准确的6D物体姿态估计是具身智能体的基础能力,但在开放世界环境中仍极具挑战。现有方法常依赖封闭集假设或几何无关的回归方案,限制了其泛化性、稳定性与机器人系统的实时应用。我们提出OMNI-PoseX,一种视觉基础模型,引入新颖的网络架构,统一开放词汇感知与SO(3)-感知的反射流匹配姿态预测器。该架构将对象级理解与几何一致的旋转推断解耦,并采用轻量级多模态融合策略,以紧凑的语义嵌入条件化敏感于旋转的几何特征,实现高效稳定的6D姿态估计。为增强鲁棒性与泛化能力,模型在大规模6D姿态数据集上训练,涵盖广泛物体多样性、视角变化与场景复杂性,构建可扩展的开放世界姿态主干。在基准姿态估计、消融实验、零样本泛化及系统级机器人抓取集成中的全面评估表明,OMNI-PoseX在保持最先进姿态精度的同时,具备实时效率,且输出几何一致的预测,可可靠抓取多种此前未见物体。
原文摘要 · Abstract (English)
Accurate 6D object pose estimation is a fundamental capability for embodied agents, yet remains highly challenging in open-world environments. Many existing methods often rely on closed-set assumptions or geometry-agnostic regression schemes, limiting their generalization, stability, and real-time applicability in robotic systems. We present OMNI-PoseX, a vision foundation model that introduces a novel network architecture unifying open-vocabulary perception with an SO(3)-aware reflected flow matching pose predictor. The architecture decouples object-level understanding from geometry-consistent rotation inference, and employs a lightweight multi-modal fusion strategy that conditions rotation-sensitive geometric features on compact semantic embeddings, enabling efficient and stable 6D pose estimation. To enhance robustness and generalization, the model is trained on large-scale 6D pose datasets, leveraging broad object diversity, viewpoint variation, and scene complexity to build a scalable open-world pose backbone. Comprehensive evaluations across benchmark pose estimation, ablation studies, zero-shot generalization, and system-level robotic grasping integration demonstrate the effectiveness of OMNI-PoseX. The OMNI-PoseX achieves SOTA pose accuracy and real-time efficiency, while delivering geometrically consistent predictions that enable reliable grasping of diverse, previously unseen objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。