arXiv:2409.16830cs.ROcs.AI2024-09被引 4

用离线强化学习规划机器人信息采集路径,无需真实交互即可高效安全决策。

OffRIPP: Offline RL-based Informative Path Planning

  • 基于离线强化学习,从预收集数据中学习路径规划策略。
  • 在仿真与真实实验中均超越基线方法,信息获取效率更高。
  • 适合需要安全、低成本试错的机器人环境感知任务。

信息采集路径规划(IPP)是机器人领域的重要任务,要求智能体设计路径以在资源约束下获取目标环境中的有价值信息。强化学习(RL)已被证明在该任务中有效,但其依赖环境交互,实际应用中存在风险且成本高昂。为此,本文提出一种基于离线强化学习的IPP框架,在训练阶段无需实时交互,兼具安全性与成本效益;同时在执行阶段保持强化学习的高性能与快速计算优势。该框架采用批处理约束强化学习,有效缓解外推误差,使智能体可从任意算法生成的预收集数据集中学习。通过大量仿真与真实实验验证,结果表明该框架优于现有基线方法,充分证明了其有效性。

原文摘要 · Abstract (English)

Informative path planning (IPP) is a crucial task in robotics, where agents must design paths to gather valuable information about a target environment while adhering to resource constraints. Reinforcement learning (RL) has been shown to be effective for IPP, however, it requires environment interactions, which are risky and expensive in practice. To address this problem, we propose an offline RL-based IPP framework that optimizes information gain without requiring real-time interaction during training, offering safety and cost-efficiency by avoiding interaction, as well as superior performance and fast computation during execution -- key advantages of RL. Our framework leverages batch-constrained reinforcement learning to mitigate extrapolation errors, enabling the agent to learn from pre-collected datasets generated by arbitrary algorithms. We validate the framework through extensive simulations and real-world experiments. The numerical results show that our framework outperforms the baselines, demonstrating the effectiveness of the proposed approach.

强化学习路径规划机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。