用多目标强化学习让自动驾驶实时适配用户驾驶风格。
Multi-Objective Reinforcement Learning for Adaptable Personalized Autonomous Driving
- 通过连续权重向量动态调节效率、舒适度等驾驶目标。
- 无需重训练即可在仿真中实现风格自适应,且不降低安全与路线完成率。
- 适合追求个性化体验的自动驾驶研发人员参考。
人类驾驶员具有不同的驾驶偏好。使自动驾驶车辆适应这些偏好对建立用户信任和满意度至关重要。然而,现有端到端驾驶方法通常依赖预设驾驶风格或需要持续用户反馈进行调整,难以支持动态、情境相关的偏好变化。本文提出一种基于偏好驱动优化的多目标强化学习(MORL)方法,用于端到端自动驾驶,实现运行时对驾驶风格偏好的自适应。偏好以连续权重向量编码,用于调节效率、舒适性、速度和激进程度等可解释的驾驶目标,无需策略重训练。所提出的单策略代理在复杂混合交通场景中集成视觉感知,并在CARLA模拟器中多样城市环境中进行评估。实验结果表明,该代理能根据偏好变化动态调整驾驶行为,同时保持碰撞规避和路线完成性能。
原文摘要 · Abstract (English)
Human drivers exhibit individual preferences regarding driving style. Adapting autonomous vehicles to these preferences is essential for user trust and satisfaction. However, existing end-to-end driving approaches often rely on predefined driving styles or require continuous user feedback for adaptation, limiting their ability to support dynamic, context-dependent preferences. We propose a novel approach using multi-objective reinforcement learning (MORL) with preference-driven optimization for end-to-end autonomous driving that enables runtime adaptation to driving style preferences. Preferences are encoded as continuous weight vectors to modulate behavior along interpretable style objectives$\unicode{x2013}$including efficiency, comfort, speed, and aggressiveness$\unicode{x2013}$without requiring policy retraining. Our single-policy agent integrates vision-based perception in complex mixed-traffic scenarios and is evaluated in diverse urban environments using the CARLA simulator. Experimental results demonstrate that the agent dynamically adapts its driving behavior according to changing preferences while maintaining performance in terms of collision avoidance and route completion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。