用时序图注意力与分层融合提升无人车在人群中的导航能力
DRL-TH: Jointly Utilizing Temporal Graph Attention and Hierarchical Fusion for UGV Navigation in Crowded Environments
- 引入时序图注意力网络,捕捉连续帧间关联
- 分层图池化实现多模态特征动态融合,提升适应性
- 在真实无人车上验证,适用于复杂人流场景
深度强化学习(DRL)在拥挤环境中无人地面车辆(UGV)自主导航与避障方面展现出潜力。现有方法多依赖单帧观测并采用简单拼接进行多模态融合,难以捕捉时序上下文,限制了动态适应能力。为此,我们提出一种基于DRL的导航框架DRL-TH,利用时序图注意力和分层图池化,整合历史观测并自适应融合多模态信息。具体地,设计时序引导图注意力网络(TG-GAT),将时序权重融入注意力得分,隐式估计场景演化;构建图层次抽象模块(GHAM),通过分层池化与可学习加权融合,动态整合RGB与LiDAR特征,实现多尺度平衡表征。大量实验表明,DRL-TH在多种拥挤环境中优于现有方法。我们还在真实UGV上部署控制策略,验证其在现实场景中的良好表现。
原文摘要 · Abstract (English)
Deep reinforcement learning (DRL) methods have demonstrated potential for autonomous navigation and obstacle avoidance of unmanned ground vehicles (UGVs) in crowded environments. Most existing approaches rely on single-frame observation and employ simple concatenation for multi-modal fusion, which limits their ability to capture temporal context and hinders dynamic adaptability. To address these challenges, we propose a DRL-based navigation framework, DRL-TH, which leverages temporal graph attention and hierarchical graph pooling to integrate historical observations and adaptively fuse multi-modal information. Specifically, we introduce a temporal-guided graph attention network (TG-GAT) that incorporates temporal weights into attention scores to capture correlations between consecutive frames, thereby enabling the implicit estimation of scene evolution. In addition, we design a graph hierarchical abstraction module (GHAM) that applies hierarchical pooling and learnable weighted fusion to dynamically integrate RGB and LiDAR features, achieving balanced representation across multiple scales. Extensive experiments demonstrate that our DRL-TH outperforms existing methods in various crowded environments. We also implemented DRL-TH control policy on a real UGV and showed that it performed well in real world scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。