arXiv:2603.24138cs.LGcs.SY2026-03中稿 · ECC 2026被引 1

用数值数据和人类偏好联合优化自动驾驶轨迹,减少对真人测试的依赖。

Efficient Controller Learning from Human Preferences and Numerical Data Via Multi-Modal Surrogate Models

  • 融合低精度数值数据与高精度人类偏好,构建多模态优化框架。
  • 在自动驾驶场景中,减少80%以上的人类参与实验次数。
  • 适合需要个性化交互的智能系统设计,如自动驾驶与人机协同。

手动调整控制策略以满足高层目标通常耗时较长。贝叶斯优化提供了一种高效的数据驱动框架,通过量化目标函数评估来自动化该过程。然而,许多系统(尤其是涉及人类的)需基于主观标准进行优化。偏好型贝叶斯优化通过学习成对比较而非定量测量来解决此问题,但仅依赖偏好数据效率较低。本文提出一种多保真度、多模态贝叶斯优化框架,整合低保真度数值数据与高保真度人类偏好。该方法采用具有分层自回归结构与非分层协相关结构的高斯过程代理模型,实现对多模态数据的高效学习。我们在自动驾驶轨迹规划器的调优中验证了该框架,结果表明:结合数值数据与偏好数据可显著减少对人类决策者的实验需求,同时有效适配个体驾驶风格。

原文摘要 · Abstract (English)

Tuning control policies manually to meet high-level objectives is often time-consuming. Bayesian optimization provides a data-efficient framework for automating this process using numerical evaluations of an objective function. However, many systems, particularly those involving humans, require optimization based on subjective criteria. Preferential Bayesian optimization addresses this by learning from pairwise comparisons instead of quantitative measurements, but relying solely on preference data can be inefficient. We propose a multi-fidelity, multi-modal Bayesian optimization framework that integrates low-fidelity numerical data with high-fidelity human preferences. Our approach employs Gaussian process surrogate models with both hierarchical, autoregressive and non-hierarchical, coregionalization-based structures, enabling efficient learning from mixed-modality data. We illustrate the framework by tuning an autonomous vehicle's trajectory planner, showing that combining numerical and preference data significantly reduces the need for experiments involving the human decision maker while effectively adapting driving style to individual preferences.

贝叶斯优化人机交互自动驾驶多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。