arXiv:2606.05248cs.RO2026-06

用符号规划与强化学习结合,让机器人逆向完成操作任务。

Inverse Manipulation through Symbolic Planning and Residual Operator Learning

论文配图:Inverse Manipulation through Symbolic Planning and Residual Operator Learning
图 1 · 摘自论文原文
  • 从示范中提取符号操作符,构建逆向目标
  • 符号逆向+强化学习补足误差,精准恢复物体位置
  • 适合需精确逆向操作的机器人任务研究者

逆向机器人操作不仅需要逆转符号状态转移或回溯运动轨迹。在连续交互动态下,符号逆向计划常无法完全还原正向执行的效果。本文提出一种混合框架,通过软几何谓词自动从示范中提取类似STRIPS的操作符,并为每个操作构建逆向恢复目标:保留前提条件、恢复删除效应、否定添加效应。任务规划器首先尝试用可用动作原语满足该目标;未解决的符号谓词则触发残差操作学习问题,通过强化学习(如软演员-评论家)求解。我们在ManiSkill3 PushCube任务上验证了该方法。对于正向推送技能,符号逆向实现粗略的拾取-放置恢复,而残差软演员-评论家策略进一步精调立方体姿态以满足剩余逆向谓词。结果表明,基于谓词的残差控制可将近似符号逆向转化为物理上成立的逆向技能。

原文摘要 · Abstract (English)

Inverting a robotic task requires more than reversing symbolic state transitions or rewinding motor trajectories. In robot manipulation tasks, symbolic inverse plans often fail to fully restore the effects of forward executions under continuous interaction dynamics. We present a hybrid framework for inverse manipulation that derives inverse-skill objectives from STRIPS-like operators automatically extracted from demonstrations through soft geometric predicates. For each extracted operator, we construct an inverse restoration objective that preserves preconditions, restores delete effects, and negates add effects. A task planner first attempts to satisfy this objective using available action primitives. Unresolved symbolic predicates then induce a residual operator learning problem solved through Reinforcement Learning (RL). We evaluate the framework on the ManiSkill3 PushCube task. For a forward pushing skill, the symbolic inverse performs a coarse pick-and-place restoration, while a residual Soft Actor-Critic policy refines the cube pose to satisfy the remaining inverse predicates. Our results show that predicate-derived residual control can turn an approximate symbolic inverse into a physically grounded inverse skill.

机器人逆向操作符号规划强化学习逆向技能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。