arXiv:2501.18848cs.RO2025-01中稿 · ICRA被引 2

让机器人灵活理解符号指令,避免漏检。

Reinforcement Learning of Flexible Policies for Symbolic Instructions with Adjustable Mapping Specifications

  • 将符号指令与映射规范分离建模,支持多状态匹配同一符号。
  • 在3D仿真中优于对比方法,实现复杂任务的高效学习。
  • 适合需要多角度评估的工业检测场景,如设备巡检。

符号化任务表示是编码人类指令和领域知识的强大工具。此类指令通过强化学习(RL)引导机器人完成多样化目标并满足约束条件。现有方法通常基于环境状态到符号的固定映射,但在需多视角评估设备状态的检测任务中,机器人必须从不同状态触发相同符号。为应对这一灵活性需求,本文提出将符号及其映射规范在强化学习策略中分别建模,使策略学习符号指令与映射规范的组合。为此,我们设计了符号指令可调映射规范(SIAMS)方法,采用线性时序逻辑(LTL)表示符号指令,便于集成至强化学习框架。该方法通过(1)感知规范的状态调制,将映射规范差异嵌入状态特征;(2)基于符号数量的任务课程,按学习进度逐步提供任务。在离散与连续动作空间的3D仿真中评估表明,本方法优于上下文感知的多任务强化学习基线。

原文摘要 · Abstract (English)

Symbolic task representation is a powerful tool for encoding human instructions and domain knowledge. Such instructions guide robots to accomplish diverse objectives and meet constraints through reinforcement learning (RL). Most existing methods are based on fixed mappings from environmental states to symbols. However, in inspection tasks, where equipment conditions must be evaluated from multiple perspectives to avoid errors of oversight, robots must fulfill the same symbol from different states. To help robots respond to flexible symbol mapping, we propose representing symbols and their mapping specifications separately within an RL policy. This approach imposes on RL policy to learn combinations of symbolic instructions and mapping specifications, requiring an efficient learning framework. To cope with this issue, we introduce an approach for learning flexible policies called Symbolic Instructions with Adjustable Mapping Specifications (SIAMS). This paper represents symbolic instructions using linear temporal logic (LTL), a formal language that can be easily integrated into RL. Our method addresses the diversified completion patterns of instructions by (1) a specification-aware state modulation, which embeds differences in mapping specifications in state features, and (2) a symbol-number-based task curriculum, which gradually provides tasks according to the learning's progress. Evaluations in 3D simulations with discrete and continuous action spaces demonstrate that our method outperforms context-aware multitask RL comparisons.

强化学习符号指令灵活映射机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。