arXiv:2512.06017cs.ROeess.IV2025-12中稿 · CVIS 2025

用现成大模型无训练估计机器人关节角度

Training-Free Robot Pose Estimation using Off-the-Shelf Foundational Models

  • 直接使用现成视觉语言模型,无需额外训练
  • 在合成与真实图像上验证了当前大模型的性能基准
  • 发现测试时缩放或参数缩放无法提升预测精度

从视觉输入中估计机器人臂的位姿是一项具有挑战性的任务。随着机器人臂在工业和家庭场景中的广泛应用,可靠的关节角度估计能提升安全性和性能,也可用于验证并进一步训练机器人策略。本文提出直接使用前沿视觉语言模型(VLMs)作为“即插即用”工具,仅通过单张目标图像即可估计机器人臂的关节角度。通过对合成与真实图像数据对评估前沿VLMs,本文建立了当前基础大模型(FLMs)能达到的性能基准。此外,实验结果表明,仅通过测试时缩放或参数缩放,并不能带来更好的关节角度预测效果。

原文摘要 · Abstract (English)

Pose estimation of a robot arm from visual inputs is a challenging task. However, with the increasing adoption of robot arms for both industrial and residential use cases, reliable joint angle estimation can offer improved safety and performance guarantees, and also be used as a verifier to further train robot policies. This paper introduces using frontier vision-language models (VLMs) as an ``off-the-shelf" tool to estimate a robot arm's joint angles from a single target image. By evaluating frontier VLMs on both synthetic and real-world image-data pairs, this paper establishes a performance baseline attained by current FLMs. In addition, this paper presents empirical results suggesting that test time scaling or parameter scaling alone does not lead to improved joint angle predictions.

机器人位姿估计视觉语言模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。