arXiv:2504.15329cs.GRcs.CV2025-04被引 1

Vision6D让研究人员在2D场景中交互式标注3D物体的6自由度位姿。

Vision6D: 3D-to-2D Interactive Visualization and Annotation Tool for 6D Pose Estimation

  • 通过3D-2D交互界面,直观可视化并调整物体在真实场景中的位姿。
  • 在Linemod和HANDAL数据集上验证,手动标注精度显著优于默认真值。
  • 适合需要高精度位姿标注的研究者,尤其适用于相机内外参未知场景。

6D位姿估计在机器人与物理物体精确交互的任务中日益重要。本文提出首个支持用户在2D真实场景中交互式可视化与操作3D物体的工具Vision6D,包含全面的用户研究。该系统通过视觉提示和空间关系,实现鲁棒的6D相机位姿标注,仅需相机内参即可准确标注物体在不同环境中的位置与姿态。这一能力为跨领域先进位姿估计模型的训练与发展奠定基础。通过在主流开源数据集Linemod和HANDAL上的对比实验,验证了手动标注相比默认真值具有更高精度;用户研究也表明,借助直观的3D界面,用户能高效生成准确位姿标注。该工具旨在弥合2D场景投影与3D场景之间的差距,为研究人员提供有效的6D位姿标注解决方案。软件已开源,可访问https://github.com/InteractiveGL/vision6D获取。

原文摘要 · Abstract (English)

Accurate 6D pose estimation has gained more attention over the years for robotics-assisted tasks that require precise interaction with physical objects. This paper presents an interactive 3D-to-2D visualization and annotation tool to support the 6D pose estimation research community. To the best of our knowledge, the proposed work is the first tool that allows users to visualize and manipulate 3D objects interactively on a 2D real-world scene, along with a comprehensive user study. This system supports robust 6D camera pose annotation by providing both visual cues and spatial relationships to determine object position and orientation in various environments. The annotation feature in Vision6D is particularly helpful in scenarios where the transformation matrix between the camera and world objects is unknown, as it enables accurate annotation of these objects' poses using only the camera intrinsic matrix. This capability serves as a foundational step in developing and training advanced pose estimation models across various domains. We evaluate Vision6D's effectiveness by utilizing widely-used open-source pose estimation datasets Linemod and HANDAL through comparisons between the default ground-truth camera poses with manual annotations. A user study was performed to show that Vision6D generates accurate pose annotations via visual cues in an intuitive 3D user interface. This approach aims to bridge the gap between 2D scene projections and 3D scenes, offering an effective way for researchers and developers to solve 6D pose annotation related problems. The software is open-source and publicly available at https://github.com/InteractiveGL/vision6D.

6D位姿交互标注三维可视化机器人感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。