arXiv:2509.23575cs.RO2025-09被引 4

通过语言对齐3D关键点实现机器人操作的泛化,仅用少量数据即可精准执行新指令。

Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints

  • 分层策略先定位兴趣区域,再精细控制动作,提升样本效率。
  • 在模拟与真实机器人上均实现12%成功率提升,训练轨迹仅为五分之一。
  • 适合需要少样本学习和跨场景泛化的机器人任务开发者使用。

分层粗到精策略通过先预测兴趣区域再指导精细动作,显著提升机器人3D操作的样本效率与精度。但即便结合预训练模型,仍存在泛化能力不足的问题。为此,我们提出语言对齐的粗到精操作框架(CLAP),包含三项核心组件:任务分解、基于视觉语言模型的3D关键点预测微调,以及3D感知表示。在仿真与真实机器人上广泛实验表明其卓越泛化能力。在专为评估泛化性设计的GemBench基准上,相比最先进方法平均成功率高出12%,且仅需1/5的训练轨迹。真实实验中,仅用10次示范训练,该策略即可成功泛化至新指令与新环境。

原文摘要 · Abstract (English)

Hierarchical coarse-to-fine policy, where a coarse branch predicts a region of interest to guide a fine-grained action predictor, has demonstrated significant potential in robotic 3D manipulation tasks by especially enhancing sample efficiency and enabling more precise manipulation. However, even augmented with pre-trained models, these hierarchical policies still suffer from generalization issues. To enhance generalization to novel instructions and environment variations, we propose Coarse-to-fine Language-Aligned manipulation Policy (CLAP), a framework that integrates three key components: 1) task decomposition, 2) VLM fine-tuning for 3D keypoint prediction, and 3) 3D-aware representation. Through comprehensive experiments in simulation and on a real robot, we demonstrate its superior generalization capability. Specifically, on GemBench, a benchmark designed for evaluating generalization, our approach achieves a 12\% higher average success rate than the SOTA method while using only 1/5 of the training trajectories. In real-world experiments, our policy, trained on only 10 demonstrations, successfully generalizes to novel instructions and environments.

机器人操作语言对齐3D关键点少样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。