arXiv:2412.06636math.OCcs.RO2024-12被引 1

用自由能门控机制组合简单动作,让机器人在障碍环境中自主导航。

Free-Gate: Planning, Control And Policy Composition via Free Energy Gating

  • 通过最小化自由能的门控机制组合控制原语,实现智能决策。
  • 即使环境非线性、随机且时变,优化问题仍保持凸性,保证解的可靠性。
  • 适用于仅有基础动作模块的机器人系统,适合强化学习与具身智能研究者。

本文研究如何最优地组合一组基本控制原语以完成规划与控制任务。为此,提出一种基于自由能计算的规划与控制模型——Free-Gate。该模型通过门控机制组合控制原语,以最小化变分自由能。该组合问题被建模为有限时域最优控制问题,并证明即使代价函数在状态/动作上非凸,且环境为非线性、随机、非平稳的情况下,该问题仍保持凸性。我们开发了一种算法以计算最优原语组合,并通过仿真和硬件实验验证了其有效性。实验表明,在仅具备简单电机原语(单独无法完成任务)的条件下,机器人仍能成功导航至目标位置。

原文摘要 · Abstract (English)

We consider the problem of optimally composing a set of primitives to tackle planning and control tasks. To address this problem, we introduce a free energy computational model for planning and control via policy composition: Free-Gate. Within Free-Gate, control primitives are combined via a gating mechanism that minimizes variational free energy. This composition problem is formulated as a finite-horizon optimal control problem, which we prove remains convex even when the cost is not convex in states/actions and the environment is nonlinear, stochastic and non-stationary. We develop an algorithm that computes the optimal primitives composition and demonstrate its effectiveness via in-silico and hardware experiments on an application involving robot navigation in an environment with obstacles. The experiments highlight that Free-Gate enables the robot to navigate to the destination despite only having available simple motor primitives that, individually, could not fulfill the task.

强化学习机器人控制最优控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。