arXiv:2607.16312cs.CVcs.RO2026-07

无需训练即可精准识别物体姿态,让机器人抓取更灵活。

xperception -- Making Robotic Grasping Easier

论文配图:xperception -- Making Robotic Grasping Easier
图 1 · 摘自论文原文
  • 直接用标准CAD模型+大模型语义特征,实现零样本6D姿态估计。
  • 在严重遮挡下仍保持毫米级精度,支持工业级边缘部署。
  • 已通过技术成熟度6级验证,适合高混小批量制造场景。

高混小批量制造需求推动机器人操作灵活性提升。但传统视觉系统在引入新物体时需大量数据收集与模型重训,成为瓶颈。为此,我们提出xperception,一种零样本6D姿态估计技术,无需针对特定物体微调或繁琐标注。该方法直接利用通用CAD模型,并融合基础模型(如DINOv2、GeDi)的丰富语义特征,实现毫米级精度的姿态估计。xperception在箱体拣选等工业任务中表现出对严重遮挡的鲁棒性,且可部署于NVIDIA Jetson Thor等工业边缘硬件。经验证,其核心技术基于获2024年BOP挑战赛冠军的FreeZe算法,已在技术成熟度6级下完成测试,为非结构化高混小批量制造提供可扩展、即插即用的机器人自动化方案。

原文摘要 · Abstract (English)

The transition toward high-mix low-volume manufacturing demands flexibility in robotic manipulation. However, conventional vision systems remain a bottleneck, requiring extensive data collection and model retraining whenever a new object is introduced to the production line. To overcome this rigidity, we present xperception, a zero-shot 6D pose estimation technology that eliminates the need for object-specific fine-tuning and laborious data annotation. By directly utilizing typical CAD models and integrating the rich semantic features of foundation models (e.g. DINOv2, GeDi), xperception achieves millimeter-accurate 6D pose estimation. xperception showed robustness against severe occlusions in industrial tasks like bin picking and is engineered for deployment on industrial edge hardware, such as NVIDIA Jetson Thor. Validated at a TRL of 6, the core methodology behind xperception is based on the FreeZe algorithm, which won the international BOP Challenge 2024, paving the way for scalable, plug-and-play robotic automation in unstructured high-mix low-volume manufacturing industries.

机器人抓取6D姿态估计零样本学习工业自动化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。