arXiv:2511.11052cs.RO2025-11被引 3

让机器人智能切换抓握与推挤动作,完成复杂操作。

AdaptPNP: Integrating Prehensile and Non-Prehensile Skills for Adaptive Robotic Manipulation

  • 用视觉语言模型理解任务和环境,规划抓握与非抓握动作序列。
  • 通过数字孪生预测物体位置,提前演练操作流程。
  • 支持真实机器人在线调整计划,适合需要灵活操作的场景。

非抓握(NP)操作指机器人在不形成稳定抓握的情况下改变物体状态(如推、戳、滑),当抓握不可行或不足时,可显著拓展机器人的操作能力。然而,实现跨任务、对象与环境的统一框架,无缝整合抓握(P)与非抓握(NP)动作仍具挑战:机器人需判断何时使用NP技能,选择合适动作原语,并将P与NP策略组合为稳健的多步计划。本文提出AdaptPNP,一个由视觉语言模型(VLM)驱动的任务与运动规划框架,系统性地选择并融合P与NP技能以达成多样化的操作目标。该方法利用VLM解析视觉场景与文本任务描述,生成包含P与NP动作顺序与协同关系的高层计划骨架。基于数字孪生的物体中心中间层预测期望物体位姿,实现操作序列的主动心理预演。最后,控制模块合成底层机器人指令,结合连续执行反馈,通过VLM实现在线任务计划优化与自适应重规划。我们在仿真与真实世界环境中评估了AdaptPNP在典型混合式P&NP操作任务中的表现,结果表明,混合式P&NP操作是迈向通用化、类人级机器人操作的关键一步。

原文摘要 · Abstract (English)

Non-prehensile (NP) manipulation, in which robots alter object states without forming stable grasps (for example, pushing, poking, or sliding), significantly broadens robotic manipulation capabilities when grasping is infeasible or insufficient. However, enabling a unified framework that generalizes across different tasks, objects, and environments while seamlessly integrating non-prehensile and prehensile (P) actions remains challenging: robots must determine when to invoke NP skills, select the appropriate primitive for each context, and compose P and NP strategies into robust, multi-step plans. We introduce ApaptPNP, a vision-language model (VLM)-empowered task and motion planning framework that systematically selects and combines P and NP skills to accomplish diverse manipulation objectives. Our approach leverages a VLM to interpret visual scene observations and textual task descriptions, generating a high-level plan skeleton that prescribes the sequence and coordination of P and NP actions. A digital-twin based object-centric intermediate layer predicts desired object poses, enabling proactive mental rehearsal of manipulation sequences. Finally, a control module synthesizes low-level robot commands, with continuous execution feedback enabling online task plan refinement and adaptive replanning through the VLM. We evaluate ApaptPNP across representative P&NP hybrid manipulation tasks in both simulation and real-world environments. These results underscore the potential of hybrid P&NP manipulation as a crucial step toward general-purpose, human-level robotic manipulation capabilities. Project Website: https://adaptpnp.github.io/

机器人操作视觉语言模型混合操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。