从示范中逆向学习时间约束逻辑,让机器人复现有时间顺序要求的行为。
ILCL: Inverse Logic-Constraint Learning from Temporally Constrained Demonstrations
- 用遗传算法挖掘带参数的时序逻辑表达式,无需预设模板。
- 在4个任务上超越现有方法,能有效学习并迁移时间约束。
- 适合需要精确时序行为控制的机器人应用,如装配操作。
我们旨在解决从示范中学习时间约束以复现类似示范的逻辑约束行为的问题。由于可能的规格说明空间呈组合爆炸,且非马尔可夫约束定义不明确,学习逻辑约束极具挑战。为此,我们提出一种新型时间约束学习方法——逆向逻辑约束学习(ILCL)。该方法将ILCL建模为两个玩家的零和博弈:1)基于遗传算法的时间逻辑挖掘(GA-TL-Mining),2)逻辑约束强化学习(Logic-CRL)。GA-TL-Mining无需预设模板即可高效构建带参数的截断线性时序逻辑(TLTL)语法树。随后,Logic-CRL通过一种新颖的约束重分配机制,在构造出的TLTL约束下寻找最大化任务奖励的策略。评估表明,ILCL在四个时间约束任务上优于当前最优基线方法,并成功实现到真实世界浅孔插销任务的迁移。
原文摘要 · Abstract (English)
We aim to solve the problem of temporal-constraint learning from demonstrations to reproduce demonstration-like logic-constrained behaviors. Learning logic constraints is challenging due to the combinatorially large space of possible specifications and the ill-posed nature of non-Markovian constraints. To figure it out, we introduce a novel temporal-constraint learning method, which we call inverse logic-constraint learning (ILCL). Our method frames ICL as a two-player zero-sum game between 1) a genetic algorithm-based temporal-logic mining (GA-TL-Mining) and 2) logic-constrained reinforcement learning (Logic-CRL). GA-TL-Mining efficiently constructs syntax trees for parameterized truncated linear temporal logic (TLTL) without predefined templates. Subsequently, Logic-CRL finds a policy that maximizes task rewards under the constructed TLTL constraints via a novel constraint redistribution scheme. Our evaluations show ILCL outperforms state-of-the-art baselines in learning and transferring TL constraints on four temporally constrained tasks. We also demonstrate successful transfer to real-world peg-in-shallow-hole tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。