arXiv:2503.07111cs.ROcs.CL2025-03被引 2

无需姿态估计,直接用图像映射关节角,实现零样本泛化控制。

PoseLess: Depth-Free Vision-to-Joint Control via Direct Image Mapping with VLM

  • 通过随机关节配置生成合成数据,直接从图像映射到关节角
  • 在真实场景中实现高精度关节角预测,无需人工标注数据
  • 支持机器人手到人手的跨形态迁移,适合无深度信息场景

本文提出PoseLess,一种新型机器人手部控制框架,通过将2D图像直接映射到关节角,无需显式姿态估计。该方法利用随机关节配置生成的合成数据,在零样本条件下实现真实场景泛化,并可实现从机器人手到人手的跨形态迁移。通过投影视觉输入并采用基于Transformer的解码器,PoseLess在深度模糊和数据稀缺环境下仍保持鲁棒、低延迟控制。实验表明,其关节角预测精度具有竞争力,且不依赖任何人工标注数据。

原文摘要 · Abstract (English)

This paper introduces PoseLess, a novel framework for robot hand control that eliminates the need for explicit pose estimation by directly mapping 2D images to joint angles using projected representations. Our approach leverages synthetic training data generated through randomized joint configurations, enabling zero-shot generalization to real-world scenarios and cross-morphology transfer from robotic to human hands. By projecting visual inputs and employing a transformer-based decoder, PoseLess achieves robust, low-latency control while addressing challenges such as depth ambiguity and data scarcity. Experimental results demonstrate competitive performance in joint angle prediction accuracy without relying on any human-labelled dataset.

手部控制图像映射零样本跨形态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。