用奖励佩特里网解析时序行为树,让机器人长程任务更易学。
A Reward-Petri-Net Interpretation of Temporal Behavior Trees

- 将时序行为树转为带奖励的佩特里网,自动生成学习奖励
- 在复杂环境中使强化学习成功收敛,样本效率提升明显
- 适合需要层次化与时间约束的任务,如机器人控制
本文提出将时序行为树(TBT)解释为奖励佩特里网(RPN),用于强化学习(RL)。设计具有层次结构和时间约束的复杂长周期机器人任务的奖励函数极为困难。TBT通过在叶节点引入线性时序逻辑(LTL)来扩展传统行为树(BT),不仅表达任务的结构(如顺序、选择、并行),还编码时间约束。本工作将TBT转化为佩特里网(PN),并基于其结构自动分配奖励,形成RPN。在一系列逐步增加难度的环境中,我们验证了基于TBT的奖励能使传统强化学习无法完成的任务得以学习,显著提升样本效率,并实现对学习过程的灵活、直观控制。通过不同奖励分布策略与TBT结构的组合,展示了其对学习效果的显著影响。
原文摘要 · Abstract (English)
This paper introduces an interpretation of Temporal Behavior Trees (TBTs) as Reward-Petri-Nets (RPNs) for reinforcement learning (RL). Designing reward functions for complex, long-horizon robotic tasks is notoriously difficult, especially when tasks have hierarchical structure and temporal constraints. TBTs extend conventional behavior trees (BTs) used in robotic applications by incorporating temporal properties into their leaf nodes. This allows TBTs to represents not only the behavioral task structure defined by BT operators such as Sequence, Fallback, and Parallel, but also the task's temporal constraints. In this work, the constraints are specified in the leaf nodes using Linear Temporal Logic. In order to inform RL rewards using TBTs, we provide a translation from TBT into a Petri Net (PN) and show how rewards can be automatically assigned based on the TBT's structure, resulting in a RPN. In a series of increasingly challenging environments, we demonstrate how TBT-based rewards enable learning where vanilla RL fails, improve sample efficiency, and offer flexible, intuitive control over the learning progress. We showcase the learning impact by using different reward distribution schemes and TBT structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。