arXiv:2505.23857cs.LGcs.AI2025-05

用轻量编码器提升部分可观测环境下的强化学习性能

A Convolution and Attention Based Encoder for Reinforcement Learning under Partial Observability

  • 用卷积与自注意力融合编码观察历史序列
  • 在连续控制任务上超越传统递归与Transformer模型
  • 适合需要高效推理的实时决策系统

部分可观测马尔可夫决策过程(POMDPs)因状态信息不完整,仍是强化学习的核心挑战。本文将POMDP重构为以固定长度观察历史作为增强状态的完全可观测过程。为高效编码这些历史,提出一种基于深度可分离卷积与自注意力的轻量级时序编码器,避免了循环神经网络和Transformer模型的计算开销。该方法集成到演员-评论家框架中,在部分可观测条件下的连续控制基准测试中表现更优。研究表明,轻量级时序编码能有效提升人工智能系统在不确定性下的可扩展性,推动智能体在信息不全或延迟的真实环境中实现稳健推理。

原文摘要 · Abstract (English)

Partially Observable Markov Decision Processes (POMDPs) remain a core challenge in reinforcement learning due to incomplete state information. We address this by reformulating POMDPs as fully observable processes with fixed-length observation histories as augmented states. To efficiently encode these histories, we propose a lightweight temporal encoder based on depthwise separable convolution and self-attention, avoiding the overhead of recurrent and Transformer-based models. Integrated into an actor-critic framework, our method achieves superior performance on continuous control benchmarks under partial observability. More broadly, this work shows that lightweight temporal encoding can improve the scalability of AI systems under uncertainty. It advances the development of agents capable of reasoning robustly in real-world environments where information is incomplete or delayed.

强化学习时序建模轻量设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。