用结构化强化学习让智能车高效引导其他车辆选路,提升整体通行效率。
Structured Reinforcement Learning for Bayesian Persuasion : Application to Intelligent Interactive Driving

- 设计基于贝叶斯劝说的信号策略,利用结构化强化学习加速在线学习。
- 在真实车道选择场景中,相比现有方法降低30%出行成本,提升双方收益。
- 适用于需要长期决策的智能交通系统,尤其适合远见型驾驶行为建模。
交互式驾驶中,具备实时交通数据的智能先导车辆可协调联网车辆的路径选择,为动态交通管理提供新思路。针对决策协调难题,本文采用贝叶斯劝说的战略信息揭示框架:主导方(先导车)通过选择性披露前方实时交通信息等信号,引导代理方(联网车辆)在部分可观测的序列决策中向自身目标靠拢。然而,代理方为最大化长期回报而采取远见行为,使主导方的信号策略设计面临巨大计算挑战。本文提出一种在线结构化强化学习框架,用于生成对远见代理具有说服力且计算高效的信号策略。主要贡献包括:(i) 针对近似最优响应的单调代理,提出MAPL算法,实现更快的在线学习;(ii) 识别主导方Q函数超模性的充分条件;(iii) 确定主导方信号策略具有说服力的充分条件;(iv) 提出超模Q学习(SQP),利用主导方动作价值的超模结构,合成高效且具说服力的信号策略;(v) 在真实场景下进行数值分析,验证该方法在车道选择任务中相比现有信号策略设计方法,能实现30%的成本节约,显著优化先导车与联网车辆的出行收益。
原文摘要 · Abstract (English)
Interactive driving, wherein an intelligent lead vehicle equipped with real-time traffic data coordinates route choices of connected vehicles, offers a promising approach to dynamic traffic management. To address the challenge of harmonising decisions, this paper considers the strategic information revealing framework of Bayesian persuasion. Here, the principal (lead vehicle) aims to guide the agent's (connected vehicle) partially observable sequential decision making towards its own objectives by selectively revealing information, such as real-time traffic ahead, using signals. However, the agent's farsighted response to maximize its long-term reward, renders the principal's signaling strategy design computationally challenging. We propose an online structured reinforcement learning framework to synthesize computationally efficient signaling strategy which is persuasive for a far-sighted agent. The main contributions of the paper are as follows: (i) For a monotonic agent with approximate best response, we propose MAPL, a structured policy learning algorithm for faster online learning, (ii) Identification of sufficient conditions for the supermodular structure of the Q function of the principal for a monotonic agent, (iii) Identification of sufficient conditions to ensure the persuasiveness of the principal's signaling strategy, (iv) Supermodular Q learning for Principal (SQP), which leverages the supermodular structure of principal's action value to synthesize computationally efficient signaling strategy that is persuasive for a monotonic learning agent, (v) Numerical analysis considering a real-time application of Bayesian persuasive driving for lane selection demonstrates that the proposed method is 30% cost efficient for optimising travelling rewards of both the lead and connected vehicle compared to the existing methodologies for signaling strategy design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。