用物理能量信息提升强化学习的效率与稳定性,适用于自动驾驶等安全关键场景。
Hybrid Energy-Aware Reward Shaping: A Unified Lightweight Physics-Guided Methodology for Policy Optimization
- 将已知的物理能量项直接作为奖励势能,每步计算开销为O(n)
- 在4个基准任务中提升收敛速度、策略稳定性和最终性能
- 理论保证能量优化有效,适合高安全性要求的控制任务
连续控制中的深度强化学习常面临方差大、能耗高、分布外泛化差的问题,因纯数据驱动探索忽略了可用的物理结构。本文提出混合能量感知奖励塑造(H-EARS),将先验已知的主要能量项直接编码为每步计算复杂度为O(n)的奖励势能。H-EARS将塑造势能分解为任务导向与能量基础两部分,并引入动作正则化项,主动调整优化目标以实现节能控制。建立了完整理论基础:塑造与正则化的函数独立性、正定海森矩阵下的能量梯度增强、函数逼近下的收敛性保证及近似势能误差界。在四个连续控制基准和四种基线算法上,H-EARS均实现一致性能提升。高保真车辆仿真验证其在极端道路条件下对安全关键场景的适用性。
原文摘要 · Abstract (English)
Deep reinforcement learning for continuous control often suffers from high variance, low energy efficiency, and poor generalization under distribution shift, as purely data-driven exploration ignores available physical structure. This paper proposes Hybrid Energy-Aware Reward Shaping (H-EARS), which encodes dominant energy terms -- assumed known a priori -- directly as reward potentials at O(n) per-step computation. H-EARS decomposes the shaping potential into task-oriented and energy-based components, supplemented by an action regularization term that deliberately modifies the optimization objective to enforce energy-efficient control. A complete theoretical foundation is established: functional independence of shaping and regularization, energy-based gradient enrichment under positive-definite Hessian conditions, convergence guarantees under function approximation, and approximate potential error bounds. Across four continuous control benchmarks and four baseline algorithms, H-EARS achieves consistent gains in convergence speed, policy stability, and final performance. High-fidelity vehicle simulations validate applicability in safety-critical settings under extreme road conditions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。