BOP挑战2024推动6D姿态估计从实验室走向真实场景。
BOP Challenge 2024 on Model-Based and Model-Free 6D Object Pose Estimation
- 引入无模型任务,仅凭参考视频即可完成物体姿态估计。
- 新数据集BOP-H3用高分辨率设备录制,更贴近真实使用环境。
- 新方法在未见物体上实现更高精度,适合实际部署需求。
我们介绍了BOP Challenge 2024的评估方法、数据集与结果,这是系列公开竞赛中的第六届,旨在反映6D物体姿态估计及相关任务的最新进展。2024年目标是将BOP从实验室环境转向真实场景。首先,引入了新的无模型任务,即无需提供3D物体模型,方法需仅通过提供的参考视频完成物体接入。其次,定义了更实用的6D物体检测任务,测试图像中物体身份不再作为输入。第三,推出了由高分辨率传感器和AR/VR头显录制的新数据集BOP-H3,真实还原现实场景。BOP-H3包含3D模型和接入视频,支持基于模型与无模型任务。参赛者在七个赛道中竞争。值得注意的是,2024年最佳基于模型的未见物体6D定位方法FreeZeV2.1在BOP-Classic-Core上比2023年最佳方法GenFlow高出22%准确率,且虽较慢(24.9秒/图,对比2.7秒/图),但仅比2023年最佳见物方法GPose2023低4%。更实用的方法Co-op仅需0.8秒/图,且比GenFlow高13%准确率。6D检测排名与6D定位相似,但运行时间更长。在未见物体的2D检测任务中,最佳方法MUSE相比2023年最优方法CNOS提升21%–29%,但未见物体2D检测准确率仍比见物情况低35%(对比GDet2023),表明2D检测仍是6D定位/检测未见物体的主要瓶颈。在线评测系统持续开放,地址为http://bop.felk.cvut.cz/
原文摘要 · Abstract (English)
We present the evaluation methodology, datasets and results of the BOP Challenge 2024, the 6th in a series of public competitions organized to capture the state of the art in 6D object pose estimation and related tasks. In 2024, our goal was to transition BOP from lab-like setups to real-world scenarios. First, we introduced new model-free tasks, where no 3D object models are available and methods need to onboard objects just from provided reference videos. Second, we defined a new, more practical 6D object detection task where identities of objects visible in a test image are not provided as input. Third, we introduced new BOP-H3 datasets recorded with high-resolution sensors and AR/VR headsets, closely resembling real-world scenarios. BOP-H3 include 3D models and onboarding videos to support both model-based and model-free tasks. Participants competed on seven challenge tracks. Notably, the best 2024 method for model-based 6D localization of unseen objects (FreeZeV2.1) achieves 22% higher accuracy on BOP-Classic-Core than the best 2023 method (GenFlow), and is only 4% behind the best 2023 method for seen objects (GPose2023) although being significantly slower (24.9 vs 2.7s per image). A more practical 2024 method for this task is Co-op which takes only 0.8s per image and is 13% more accurate than GenFlow. Methods have similar rankings on 6D detection as on 6D localization but higher run time. On model-based 2D detection of unseen objects, the best 2024 method (MUSE) achieves 21--29% relative improvement compared to the best 2023 method (CNOS). However, the 2D detection accuracy for unseen objects is still -35% behind the accuracy for seen objects (GDet2023), and the 2D detection stage is consequently the main bottleneck of existing pipelines for 6D localization/detection of unseen objects. The online evaluation system stays open and is available at http://bop.felk.cvut.cz/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。