arXiv:2502.09393cs.ROcs.NE2025-02被引 4

用类脑高维计算提升强化学习在未知环境中的路径规划泛化能力。

Generalizable Reinforcement Learning with Biologically Inspired Hyperdimensional Occupancy Grid Maps for Exploration and Goal-Directed Path Planning

  • 用高维符号架构替代传统占用网格地图,实现类脑感知建模。
  • 在未见环境中性能提升约47%,显著增强策略泛化性。
  • 适合需要跨场景适应的自动驾驶与机器人探索任务。

实时自主系统依赖多层计算框架完成感知、目标定位和路径规划等关键任务。传统方法使用占用网格映射(OGM)将环境离散化为带概率信息的单元格,结构清晰,便于下游算法处理。近期研究引入类脑数学框架——向量符号架构(VSA),即高维计算,用于在高维空间中实现概率性OGM,称为VSA-OGM。该方法天然兼容脉冲神经网络,可作为传统OGM的类脑替代方案。然而,大规模集成前需评估其对下游任务的影响。本研究对比VSA-OGM与传统贝叶斯希尔伯特映射(BHM)在基于强化学习的目标发现与路径规划框架中的表现,涵盖受控探索环境及受F1-Tenth挑战启发的自动驾驶场景。结果表明,VSA-OGM在单场景与多场景训练下保持相当的学习性能,且在未见环境中性能提升约47%。这证明以VSA-OGM训练的策略网络具有更强的泛化能力,支持其在多样化真实环境中的部署潜力。

原文摘要 · Abstract (English)

Real-time autonomous systems utilize multi-layer computational frameworks to perform critical tasks such as perception, goal finding, and path planning. Traditional methods implement perception using occupancy grid mapping (OGM), segmenting the environment into discretized cells with probabilistic information. This classical approach is well-established and provides a structured input for downstream processes like goal finding and path planning algorithms. Recent approaches leverage a biologically inspired mathematical framework known as vector symbolic architectures (VSA), commonly known as hyperdimensional computing, to perform probabilistic OGM in hyperdimensional space. This approach, VSA-OGM, provides native compatibility with spiking neural networks, positioning VSA-OGM as a potential neuromorphic alternative to conventional OGM. However, for large-scale integration, it is essential to assess the performance implications of VSA-OGM on downstream tasks compared to established OGM methods. This study examines the efficacy of VSA-OGM against a traditional OGM approach, Bayesian Hilbert Maps (BHM), within reinforcement learning based goal finding and path planning frameworks, across a controlled exploration environment and an autonomous driving scenario inspired by the F1-Tenth challenge. Our results demonstrate that VSA-OGM maintains comparable learning performance across single and multi-scenario training configurations while improving performance on unseen environments by approximately 47%. These findings highlight the increased generalizability of policy networks trained with VSA-OGM over BHM, reinforcing its potential for real-world deployment in diverse environments.

强化学习类脑计算路径规划泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。