arXiv:2509.08813cs.RO2025-09被引 3

无需标定板,用3D模型实现机器人手眼标定与厘米级场景重建。

Calib3R: Hand-Eye Calibration and 3D Metric-Scaled Scene Reconstruction with 3D Foundation Models

  • 基于3D基础模型从单张或多张图像提取点云,联合优化标定与重建。
  • 仅需10张以内图像即可实现亚厘米级精度,优于传统标定方法。
  • 适用于单/多相机机器人系统,输出与机器人基座对齐的度量空间场景。

机器人常依赖RGB图像完成操作任务,但可靠交互需具备度量尺度且与机器人参考系对齐的三维场景表示。这依赖于精确的手眼标定和稠密三维重建,二者通常独立处理,却均基于RGB数据中的几何对应关系。传统标定方法需使用标定板,而基于RGB的重建结果仅提供无尺度的任意坐标系下的三维几何。多相机设置进一步增加复杂性,需将数据统一到共享参考系中。本文提出Calib3R,一种无需标定板的方法,通过统一优化联合完成手眼标定与度量尺度的三维重建。Calib3R可处理机器人臂上的单/多相机配置。其基于3D基础模型从RGB图像提取点图(pointmaps),结合机器人位姿重建与机器人基座参考系对齐的缩放三维场景。在多个数据集上的实验表明,Calib3R仅需少于10张图像即可实现高精度标定,性能优于基于模式和无模式的现有方法。

原文摘要 · Abstract (English)

Robots often rely on RGB images for tasks like manipulation. However, reliable interaction typically requires a 3D scene representation that is metric-scaled and aligned with the robot reference frame. This depends on accurate hand-eye calibration and dense 3D reconstruction, tasks usually treated separately, despite both relying on geometric correspondences from RGB data. Traditional calibration techniques needs patterns, while RGB-based reconstruction yields 3D geometry with an unknown scale in an arbitrary frame. Multi-camera setups add further complexity, as data must be expressed in a shared reference frame. We present Calib3R, a patternless method that jointly performs hand-eye calibration and metric-scaled 3D reconstruction via unified optimization. Calib3R handles single- and multi-camera setups on robot arms. It builds on a 3D foundation model to extract pointmaps from RGB images, which are combined with robot poses to reconstruct a scaled 3D scene aligned with the robot base reference frame. Experiments on diverse datasets show that Calib3R achieves accurate calibration with less than 10 images, outperforming patternless and pattern-based methods.

手眼标定3D重建机器人感知3D基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。