用先验知识加速自动驾驶风格偏好学习,减少不适体验
Exploiting Prior Knowledge in Preferential Learning of Individualized Autonomous Vehicle Driving Styles

- 引入真实驾驶数据构建虚拟决策者,指导参数采样
- 相比传统方法收敛更快,减少不舒适驾驶行为次数
- 适合个性化自动驾驶系统研发与人机交互优化
自动驾驶车辆的轨迹规划常采用模型预测控制(MPC),其成本函数直接影响驾驶风格。然而,如何设计出乘客偏好的成本函数仍是挑战。本文采用偏好贝叶斯优化,通过迭代询问乘客偏好来学习成本函数。由于参数空间维度高,现有方法在有限实验下难以找到最优解,且探索过程可能让乘客感到不适。为此,我们引入先验知识,在偏好贝叶斯优化框架中构建基于真实人类驾驶数据的虚拟决策者,引导参数采样。仿真实验表明,该方法收敛速度优于现有方法,显著减少了不适宜驾驶风格的采样次数。
原文摘要 · Abstract (English)
Trajectory planning for automated vehicles commonly employs optimization over a moving horizon - Model Predictive Control - where the cost function critically influences the resulting driving style. However, finding a suitable cost function that results in a driving style preferred by passengers remains an ongoing challenge. We employ preferential Bayesian optimization to learn the cost function by iteratively querying a passenger's preference. Due to increasing dimensionality of the parameter space, preference learning approaches might struggle to find a suitable optimum with a limited number of experiments and expose the passenger to discomfort when exploring the parameter space. We address these challenges by incorporating prior knowledge into the preferential Bayesian optimization framework. Our method constructs a virtual decision maker from real-world human driving data to guide parameter sampling. In a simulation experiment, we achieve faster convergence of the prior-knowledge-informed learning procedure compared to existing preferential Bayesian optimization approaches and reduce the number of inadequate driving styles sampled.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。