用人类反馈强化学习,让自动驾驶更像不同人的开车风格。
Learning Personalized Driving Styles via Reinforcement Learning from Human Feedback
- 通过人类反馈强化学习,优化生成式轨迹模型的个性化表现。
- 在NavSim基准上达到顶尖水平,同时保持安全与可行性。
- 适合需要模仿真实驾驶习惯的自动驾驶系统研发者。
在动态环境中生成类人且自适应的行驶轨迹对自动驾驶至关重要。尽管生成模型在合成可行轨迹方面展现出潜力,但常因数据集偏差和分布偏移而难以捕捉个性化的驾驶风格差异。为此,我们提出TrajHF——一种基于人类反馈的微调框架,用于生成式轨迹模型,旨在使运动规划更贴合多样化的驾驶风格。TrajHF结合多条件去噪器与人类反馈强化学习,超越传统模仿学习,实现对人类驾驶偏好更好的对齐,同时保障安全与可行性约束。在NavSim基准测试中,TrajHF表现媲美当前最先进方法。该工作为自动驾驶中的个性化与自适应轨迹生成树立了新范式。
原文摘要 · Abstract (English)
Generating human-like and adaptive trajectories is essential for autonomous driving in dynamic environments. While generative models have shown promise in synthesizing feasible trajectories, they often fail to capture the nuanced variability of personalized driving styles due to dataset biases and distributional shifts. To address this, we introduce TrajHF, a human feedback-driven finetuning framework for generative trajectory models, designed to align motion planning with diverse driving styles. TrajHF incorporates multi-conditional denoiser and reinforcement learning with human feedback to refine multi-modal trajectory generation beyond conventional imitation learning. This enables better alignment with human driving preferences while maintaining safety and feasibility constraints. TrajHF achieves performance comparable to the state-of-the-art on NavSim benchmark. TrajHF sets a new paradigm for personalized and adaptable trajectory generation in autonomous driving.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。