arXiv:2604.15772cs.RO2026-04

用模糊逻辑动态调整奖励,让智能体更稳更快学会复杂导航。

Fuzzy Logic Theory-based Adaptive Reward Shaping for Robust Reinforcement Learning (FARS)

论文配图:Fuzzy Logic Theory-based Adaptive Reward Shaping for Robust Reinforcement Learning (FARS)
图 1 · 摘自论文原文
  • 基于模糊规则按状态动态调节奖励,融合人类经验
  • 在复杂任务中收敛更快,成功率最高提升约5%
  • 适合需要稳定学习的高维控制场景,如无人机竞速

强化学习在高维状态空间和长时序任务中常因奖励稀疏或固定而探索缓慢,易陷入局部最优。本文提出一种基于模糊逻辑的自适应奖励塑造方法(FARS),将专家知识编码为可解释的模糊规则,使奖励贡献随智能体状态动态调整,促进平稳学习并降低对超参数的敏感性。该方法在自主无人机竞速基准测试中表现优异,实现了更稳定的训练过程与一致的任务性能,尤其在难度递增的场景中,收敛速度加快,不同训练种子间的性能波动减少,成功率最高提升约5%。

原文摘要 · Abstract (English)

Reinforcement learning (RL) often struggles in real-world tasks with high-dimensional state spaces and long horizons, where sparse or fixed rewards severely slow down exploration and cause agents to get trapped in local optima. This paper presents a fuzzy logic based reward shaping method that integrates human intuition into RL reward design. By encoding expert knowledge into adaptive and interpreable terms, fuzzy rules promote stable learning and reduce sensitivity to hyperparameters. The proposed method leverages these properties to adapt reward contributions based on the agent state, enabling smoother transitions between fast motion and precise control in challenging navigation tasks. Extensive simulation results on autonomous drone racing benchmarks show stable learning behavior and consistent task performance across scenarios of increasing difficulty. The proposed method achieves faster convergence and reduced performance variability across training seeds in more challenging environments, with success rates improving by up to approximately 5 percent compared to non fuzzy reward formulations.

强化学习奖励塑造模糊逻辑无人机

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。