从视觉演示中自动学习机器人不确定环境下的符号化规划模型
PO-PDDL: Learning Symbolic POMDPs from Visual Demonstrations for Robot Planning Under Uncertainty

- 用符号化语言重构机器人执行中的隐藏状态轨迹
- 在真实长时序操作中实现比现有方法更高鲁棒性与更低规划开销
- 适合需要应对感知与动作不确定性的真实机器人场景
现实世界中的机器人任务规划需同时应对随机动作执行和部分可观测性,但为实际机器人领域构建部分可观测马尔可夫决策过程(POMDP)模型仍困难且耗时。我们提出PO-PDDL,一种保留关系结构和大语言模型友好语法的符号化POMDP形式,显式建模部分可观测性、随机性和信念。基于此,我们设计了一种示范驱动的建模流程:从真实机器人执行视频中重建隐含的符号状态轨迹,通过推断状态与视觉观测间的不一致性识别部分可观测性,并据此学习随机转移与观测模型。所得的PO-PDDL领域可跨任务复用,支持在感知与执行不确定性下进行在线信念空间规划。在真实世界长时序操作任务上的实验表明,该方法持续优于现有的PDDL与POMDP模型学习方法,在显著降低规划成本的同时实现强鲁棒性任务规划。
原文摘要 · Abstract (English)
Real-world robot task planning must operate under both stochastic action execution and partial observability, yet constructing Partially Observable Markov Decision Process (POMDP) models for real robotics domains remains difficult and labor-intensive. We introduce PO-PDDL, a symbolic formulation of POMDPs that preserves the relational structure and LLM-friendly syntax of the Planning Domain Definition Language (PDDL), while explicitly modeling partial observability, stochasticity, and beliefs. Building on this formulation, we propose a demonstration-driven pipeline for learning PO-PDDL models. The proposed method reconstructs latent symbolic state trajectories from real-robot execution videos, identifies partial observability via inconsistencies between inferred states and visual observations, and learns stochastic transition and observation models accordingly. The resulting PO-PDDL domains are reusable across tasks and enable online belief-space planning under both perception and execution uncertainty. Experiments on real-world long-horizon manipulation tasks show that our method consistently outperforms existing PDDL and POMDP model-learning approaches, achieving robust task planning under uncertainty with significantly lower planning cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。