用多目标强化学习让卡车在安全、省油和省时间间自动找平衡。
Multi-Objective Reinforcement Learning for Tactical Decision Making for Trucks in Highway Traffic
- 基于近端策略优化的多目标框架,显式学习多种驾驶策略
- 生成平滑可解释的帕累托前沿,覆盖安全、节能与效率的权衡
- 无需重训练即可切换策略,适合自动驾驶卡车场景
在高速公路上,重型车辆需在安全性、效率和运营成本之间取得平衡,这构成了一项具有挑战性的决策问题。传统单一奖励函数通过加权聚合多个目标,常掩盖其内在权衡结构。本文提出一种基于近端策略优化的多目标强化学习框架,能够在可扩展的仿真平台上学习一组明确表示这些权衡的策略。该方法生成一组帕累托最优策略,涵盖三个冲突目标:以碰撞次数和任务完成率衡量的安全性;以能耗和司机成本分别衡量的能源效率与时间效率。所得帕累托前沿平滑且可解释,支持在不同冲突目标间灵活选择驾驶行为。该框架可无缝切换不同策略而无需重新训练,为自动驾驶卡车应用提供鲁棒且自适应的决策方案。
原文摘要 · Abstract (English)
Balancing safety, efficiency, and operational costs in highway driving poses a challenging decision-making problem for heavy-duty vehicles. A central difficulty is that conventional scalar reward formulations, obtained by aggregating these competing objectives, often obscure the structure of their trade-offs. We present a Proximal Policy Optimization based multi-objective reinforcement learning framework that learns a set of policies explicitly representing these trade-offs and evaluates it on a scalable simulation platform for tactical decision making in trucks. The proposed approach learns a set of Pareto-optimal policies that capture the trade-offs among three conflicting objectives: safety, quantified in terms of collisions and successful completion; energy efficiency and time efficiency, quantified using energy cost and driver cost, respectively. The resulting Pareto frontier is smooth and interpretable, enabling flexibility in choosing driving behavior along different conflicting objectives. This framework allows seamless transitions between different driving policies without retraining, yielding a robust and adaptive decision-making strategy for autonomous trucking applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。