用逆强化学习建模驾驶行为,提升复杂交通下的预测泛化能力
Generalizable Trajectory Prediction via Inverse Reinforcement Learning with Mamba-Graph Architecture
- 通过逆强化学习推断多样奖励函数,模拟人类决策过程
- 在未见场景下泛化性能比基线高2.3倍,媲美微调效果
- 结合Mamba与图注意力网络,高效捕捉长序列与空间交互
精准的驾驶行为建模是实现安全高效轨迹预测的基础,但在复杂交通场景中仍具挑战。本文提出一种新型逆强化学习(IRL)框架,通过推断多样化的奖励函数来捕捉类人决策,实现跨场景强适应性。所学奖励函数用于最大化输出概率,融合Mamba模块实现长序列依赖高效建模,结合图注意力网络编码交通参与者间的空间交互。在城市交叉口与环形路口的综合评估表明,该方法不仅在预测精度上优于多种主流方法,且在未见场景下的泛化性能比其他基线高出2.3倍,展现出媲美微调的分布外适应能力。
原文摘要 · Abstract (English)
Accurate driving behavior modeling is fundamental to safe and efficient trajectory prediction, yet remains challenging in complex traffic scenarios. This paper presents a novel Inverse Reinforcement Learning (IRL) framework that captures human-like decision-making by inferring diverse reward functions, enabling robust cross-scenario adaptability. The learned reward function is utilized to maximize the likelihood of output by integrating Mamba blocks for efficient long-sequence dependency modeling with graph attention networks to encode spatial interactions among traffic agents. Comprehensive evaluations on urban intersections and roundabouts demonstrate that the proposed method not only outperforms various popular approaches in terms of prediction accuracy but also achieves 2.3 times higher generalization performance to unseen scenarios compared to other baselines, achieving adaptability in Out-of-Distribution settings that is competitive with fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。