arXiv:2501.09783cs.RO2025-01被引 13

用几何约束让机器人通用操作,无需训练就能适应新任务。

GeoManip: Geometric Constraints as General Interfaces for Robot Manipulation

  • 通过符号化语言解析任务中的几何关系,生成可执行的约束条件。
  • 在真实场景中实现零训练,对未见任务泛化能力强于现有方法。
  • 支持实时调整策略、从人演示和失败中学习,适合交互式机器人应用。

我们提出GeoManip框架,使通用机器人能利用物体与部件间的关系所衍生的几何约束进行操作。例如切胡萝卜需满足刀刃与胡萝卜方向垂直的几何约束。该框架通过符号语言表示这些约束,并将其转化为底层动作,弥合自然语言与机器人执行之间的差距,实现跨多样甚至未见任务、物体与场景的强泛化能力。不同于需大量训练的视觉-语言-动作模型,GeoManip采用零训练方式,依赖大基础模型:一个阶段特定的约束生成模块与一个几何解析器,用于识别约束涉及的物体部件;随后由求解器优化轨迹以满足任务描述与场景中的推断约束。此外,GeoManip支持上下文学习,具备五项人性化交互特性:实时策略调整、从人类示范学习、从失败案例学习、长时程动作规划及高效模仿学习数据采集。在仿真与真实世界场景中的大量评估表明,GeoManip达到当前最优性能,在避免昂贵模型训练的同时展现出卓越的分布外泛化能力。

原文摘要 · Abstract (English)

We present GeoManip, a framework to enable generalist robots to leverage essential conditions derived from object and part relationships, as geometric constraints, for robot manipulation. For example, cutting the carrot requires adhering to a geometric constraint: the blade of the knife should be perpendicular to the carrot's direction. By interpreting these constraints through symbolic language representations and translating them into low-level actions, GeoManip bridges the gap between natural language and robotic execution, enabling greater generalizability across diverse even unseen tasks, objects, and scenarios. Unlike vision-language-action models that require extensive training, operates training-free by utilizing large foundational models: a constraint generation module that predicts stage-specific geometric constraints and a geometry parser that identifies object parts involved in these constraints. A solver then optimizes trajectories to satisfy inferred constraints from task descriptions and the scene. Furthermore, GeoManip learns in-context and provides five appealing human-robot interaction features: on-the-fly policy adaptation, learning from human demonstrations, learning from failure cases, long-horizon action planning, and efficient data collection for imitation learning. Extensive evaluations on both simulations and real-world scenarios demonstrate GeoManip's state-of-the-art performance, with superior out-of-distribution generalization while avoiding costly model training.

机器人操作几何约束零样本泛化人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。