arXiv:2511.07717cs.ROcs.CV2025-11

用3D拓扑图联合学习2D与3D特征,减少对标注数据的依赖。

RoboTAG: End-to-end Robot Configuration Estimation via Topological Alignment Graph

  • 构建2D与3D双分支拓扑图,通过闭环一致性监督协同优化
  • 在多种机器人上验证有效,缓解真实场景数据稀缺问题
  • 适合缺乏标注数据的机器人姿态估计任务

从单目RGB图像中估计机器人位姿是机器人学与计算机视觉中的挑战。现有方法通常基于2D视觉主干网络,严重依赖标注数据进行训练,而真实场景中标签数据常稀缺,导致仿真到现实的差距。此外,这些方法将3D问题简化为2D域,忽略了3D先验信息。为此,我们提出机器人拓扑对齐图(RoboTAG),引入3D分支注入3D先验,同时实现2D与3D表示的共同演化,降低对标签的依赖。RoboTAG包含3D分支与2D分支,节点表示相机与机器人系统的状态,边捕捉变量间的依赖关系或对齐关系。图中定义闭合环路,可在分支间施加一致性监督。实验结果表明,该方法在多种机器人类型上均有效,为缓解机器人领域的数据瓶颈提供了新可能。

原文摘要 · Abstract (English)

Estimating robot pose from a monocular RGB image is a challenge in robotics and computer vision. Existing methods typically build networks on top of 2D visual backbones and depend heavily on labeled data for training, which is often scarce in real-world scenarios, causing a sim-to-real gap. Moreover, these approaches reduce the 3D-based problem to 2D domain, neglecting the 3D priors. To address these, we propose Robot Topological Alignment Graph (RoboTAG), which incorporates a 3D branch to inject 3D priors while enabling co-evolution of the 2D and 3D representations, alleviating the reliance on labels. Specifically, the RoboTAG consists of a 3D branch and a 2D branch, where nodes represent the states of the camera and robot system, and edges capture the dependencies between these variables or denote alignments between them. Closed loops are then defined in the graph, on which a consistency supervision across branches can be applied. Experimental results demonstrate that our method is effective across robot types, suggesting new possibilities of alleviating the data bottleneck in robotics.

姿态估计3D先验少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。