通过物理启发的连续时空激励,提升机器人导航安全性。
C-STEP: Continuous Space-Time Empowerment for Physics-informed Safe Reinforcement Learning of Mobile Agents
- 基于物理动态与内部状态设计安全激励机制
- 碰撞减少,避障距离缩短,行程时间几乎不变
- 适合对可解释性与安全性要求高的移动机器人系统
在复杂环境中实现安全导航仍是机器人强化学习的核心挑战。本文提出面向确定性连续空间的连续时空激励(C-STEP),一种以智能体为中心的安全度量方法。该度量可作为物理启发的内在奖励,增强正向导航奖励函数。奖励项融合智能体的内部状态(如初始速度)与前向动力学,以区分安全与高风险行为。将C-STEP与导航奖励结合,得到一个同时优化任务完成与防碰撞的内在奖励函数。数值实验表明,该方法显著减少碰撞、降低与障碍物的接近程度,且行程时间仅略有增加。总体而言,C-STEP为强化学习中的奖励塑造提供了一种可解释、物理驱动的方法,有助于提升自主移动机器人的安全性。
原文摘要 · Abstract (English)
Safe navigation in complex environments remains a central challenge for reinforcement learning (RL) in robotics. This paper introduces Continuous Space-Time Empowerment for Physics-informed (C-STEP) safe RL, a novel measure of agent-centric safety tailored to deterministic, continuous domains. This measure can be used to design physics-informed intrinsic rewards by augmenting positive navigation reward functions. The reward incorporates the agents internal states (e.g., initial velocity) and forward dynamics to differentiate safe from risky behavior. By integrating C-STEP with navigation rewards, we obtain an intrinsic reward function that jointly optimizes task completion and collision avoidance. Numerical results demonstrate fewer collisions, reduced proximity to obstacles, and only marginal increases in travel time. Overall, C-STEP offers an interpretable, physics-informed approach to reward shaping in RL, contributing to safety for agentic mobile robotic systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。