arXiv:2601.04392cs.LGcs.AI2026-01中稿 · ECC26 conference

用模糊规则提升强化学习可解释性,同时保持高效采样与稳定训练。

Enhanced-FQL($λ$), an Efficient and Interpretable RL with novel Fuzzy Eligibility Traces and Segmented Experience Replay

  • 引入模糊资格迹与分段经验回放,改进模糊Q学习的信用分配与样本效率。
  • 在倒立摆任务中样本效率更高、方差更小,性能媲美DDPG基线。
  • 适合追求可解释性与轻量级部署的连续控制场景,如机器人控制。

本文提出一种新型模糊强化学习框架 Enhanced-FQL(λ),将模糊资格迹(FET)与分段经验回放(SER)结合至模糊贝尔曼方程(FBE)下的模糊Q学习,用于连续控制任务。该方法采用可解释的模糊规则库替代复杂神经网络,通过两个关键创新维持竞争力:基于资格迹的模糊贝尔曼方程实现稳定的多步信用分配,以及基于分段的经验回放机制提升样本效率。理论分析证明了该方法在标准假设下的收敛性。在倒立摆基准测试中,Enhanced-FQL(λ)相比n步模糊TD和模糊SARSA(λ)提升了样本效率并降低了方差,且性能与测试的DDPG基线相当。结果表明,该框架是中等规模连续控制问题中兼具可解释性与计算紧凑性的有效替代方案。

原文摘要 · Abstract (English)

This paper introduces a fuzzy reinforcement learning framework, Enhanced-FQL($λ$), that integrates novel Fuzzified Eligibility Traces (FET) and Segmented Experience Replay (SER) into fuzzy Q-learning with the Fuzzified Bellman Equation (FBE) for continuous control. The proposed approach employs an interpretable fuzzy rule base instead of complex neural architectures, while maintaining competitive performance through two key innovations: a fuzzified Bellman equation with eligibility traces for stable multi-step credit assignment, and a memory-efficient segment-based experience replay mechanism for enhanced sample efficiency. Theoretical analysis proves the proposed method convergence under standard assumptions. On the Cart--Pole benchmark, Enhanced-FQL($λ$) improves sample efficiency and reduces variance relative to $n$-step fuzzy TD and fuzzy SARSA($λ$), while remaining competitive with the tested DDPG baseline. These results support the proposed framework as an interpretable and computationally compact alternative for moderate-scale continuous control problems.

强化学习模糊系统可解释性连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。