用时序逻辑约束强化学习,让机器人在不确定环境中按时完成复杂任务
Reinforcement learning with timed constraints for robotics motion planning
- 将时序逻辑公式转为时序自动机,与强化学习结合构建可训练的决策模型
- 在5×5和10×10网格及办公室场景中均成功学习满足时间约束的策略
- 适用于需要实时响应的机器人规划,尤其适合部分可观测环境
在动态且不确定的环境中,机器人系统需要满足复杂任务序列并严格遵守时间约束。度量区间时序逻辑(MITL)提供了形式化且表达力强的框架来描述此类时序要求,但将其与强化学习(RL)结合面临随机动态和部分可观测性的挑战。本文提出一种统一的基于自动机的强化学习框架,用于在马尔可夫决策过程(MDP)和部分可观测马尔可夫决策过程(POMDP)下合成满足MITL规范的策略。将MITL公式转化为时序有限确定性广义布赫自动机(Timed-LDGBA),并与底层决策过程同步,构建适用于Q-learning的产物时序模型。采用简洁而富有表现力的奖励结构,在保证时序正确性的同时支持额外性能目标。在三个仿真场景中验证:一个5×5网格世界(建模为MDP)、一个10×10网格世界(建模为POMDP)以及一个类似办公室的服务机器人场景。结果表明,该框架在随机转移下始终能学习到满足严格时间约束的策略,可扩展至更大状态空间,并在部分可观测环境中依然有效,凸显其在时间敏感且不确定性高的场景中实现可靠机器人规划的潜力。
原文摘要 · Abstract (English)
Robotic systems operating in dynamic and uncertain environments increasingly require planners that satisfy complex task sequences while adhering to strict temporal constraints. Metric Interval Temporal Logic (MITL) offers a formal and expressive framework for specifying such time-bounded requirements; however, integrating MITL with reinforcement learning (RL) remains challenging due to stochastic dynamics and partial observability. This paper presents a unified automata-based RL framework for synthesizing policies in both Markov Decision Processes (MDPs) and Partially Observable Markov Decision Processes (POMDPs) under MITL specifications. MITL formulas are translated into Timed Limit-Deterministic Generalized Büchi Automata (Timed-LDGBA) and synchronized with the underlying decision process to construct product timed models suitable for Q-learning. A simple yet expressive reward structure enforces temporal correctness while allowing additional performance objectives. The approach is validated in three simulation studies: a $5 \times 5$ grid-world formulated as an MDP, a $10 \times 10$ grid-world formulated as a POMDP, and an office-like service-robot scenario. Results demonstrate that the proposed framework consistently learns policies that satisfy strict time-bounded requirements under stochastic transitions, scales to larger state spaces, and remains effective in partially observable environments, highlighting its potential for reliable robotic planning in time-critical and uncertain settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。