arXiv:2604.04225cs.LGcs.AI2026-04

用时序行为树修复低质量演示,提升机器人学习效率。

Learning from Imperfect Demonstrations via Temporal Behavior Tree-Guided Trajectory Repair

  • 基于时序行为树识别并修复不符合规范的轨迹
  • 修复后轨迹使强化学习奖励信号更有效,提升任务成功率
  • 适合缺乏高质量示范的机器人学习场景

从演示中学习机器人控制策略是一种强大范式,但现实数据常不理想、含噪声或存在缺陷,给模仿学习和强化学习带来挑战。本文提出一个形式化框架,利用时序行为树(TBT),即结合行为树语义的信号时序逻辑扩展,修复下游策略学习前的次优轨迹。对于违反TBT规范的演示,基于模型的修复算法修正轨迹片段以满足形式化约束,生成逻辑一致且可解释的数据集。这些修复后的轨迹用于提取势函数,塑造强化学习的奖励信号,引导智能体向任务一致的状态空间区域前进,无需了解智能体的动力学模型。在离散网格世界导航及连续单/多智能体避障任务上验证了该框架的有效性,展示了其在高质量演示不可得场景下的数据高效机器人学习潜力。

原文摘要 · Abstract (English)

Learning robot control policies from demonstrations is a powerful paradigm, yet real-world data is often suboptimal, noisy, or otherwise imperfect, posing significant challenges for imitation and reinforcement learning. In this work, we present a formal framework that leverages Temporal Behavior Trees (TBT), an extension of Signal Temporal Logic (STL) with Behavior Tree semantics, to repair suboptimal trajectories prior to their use in downstream policy learning. Given demonstrations that violate a TBT specification, a model-based repair algorithm corrects trajectory segments to satisfy the formal constraints, yielding a dataset that is both logically consistent and interpretable. The repaired trajectories are then used to extract potential functions that shape the reward signal for reinforcement learning, guiding the agent toward task-consistent regions of the state space without requiring knowledge of the agent's kinematic model. We demonstrate the effectiveness of this framework on discrete grid-world navigation and continuous single and multi-agent reach-avoid tasks, highlighting its potential for data-efficient robot learning in settings where high-quality demonstrations cannot be assumed.

机器人学习轨迹修复形式化方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。