arXiv:2506.00599cs.CV2025-06被引 15

构建工业场景下6D位姿估计新基准,解决金属反光、密集堆叠难题

XYZ-IBD: Benchmarking Robust 6D Object Pose Estimation under Real-World Industrial Complexity

  • 基于75个真实工业场景,标注27.3万例金属/对称/高反光物体
  • 现有方法在该数据集上性能大幅下降,暴露工业视觉核心挑战
  • 适合机器人抓取、工业质检等复杂环境下的算法研发者使用

当前6D位姿估计基准在家庭物体上已接近饱和,但难以反映工业环境中的随机性与光学复杂性。我们提出XYZ-IBD,一个专为工业料箱抓取设计的高精度基准。该数据集包含75个多视角真实场景,约27.3万例标注实例,涵盖金属、对称及镜面物体。不同于已有数据集,其具有高密度随机堆叠与多实例模糊性,真实反映机器人操作挑战。采用多阶段半自动标注流程,确保亚毫米级标注精度,并通过误差量化方案验证质量可靠性。此外,提供大规模合成训练集,基于逼真的料箱抓取仿真渲染。对当前主流2D检测与6D位姿估计算法的基准测试显示,其性能相比家庭基准显著下降,凸显工业视觉中未解难题。该基准为复杂、高遮挡、强反射场景下的鲁棒位姿估计树立新标准。数据集与评测平台公开可获取:https://xyz-ibd.github.io。

原文摘要 · Abstract (English)

While current 6D pose estimation benchmarks have reached near-saturation on household objects, they often fail to capture the stochastic and optical complexities of industrial environments. We introduce XYZ-IBD, a high-precision benchmark for object detection and 6D pose estimation specifically designed for industrial bin-picking. XYZ-IBD addresses the domain gap by providing 75 multi-view real-world scenes containing approximately 273k annotated instances of metallic, symmetrical, and specular objects. Unlike existing datasets, our benchmark features high-density stochastic stacking and multi-instance ambiguity, reflecting authentic robotic manipulation challenges. We employ a rigorous multi-stage and semi-automatic annotation pipeline, ensuring sub-millimeter annotation accuracy. The annotations are validated through our designed error quantification scheme, securing the reliability of the annotation quality. In addition to real-world evaluation data, we provide a large-scale complementary synthetic training set that is rendered under a realistic bin-picking simulation. Benchmarking state-of-the-art (SOTA) methods for 2D detection and 6D pose estimation reveals a significant performance degradation compared to standard household benchmarks, highlighting the unsolved challenges of industrial vision. XYZ-IBD establishes a new frontier for robust pose estimation in complex, high-occlusion, and reflective scenarios. The dataset and benchmark are publicly available at https://xyz-ibd.github.io.

6D位姿估计工业视觉机器人抓取数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。