arXiv:2605.22431cs.RO2026-05

提出一种快速在线优化方法,解决未知环境下的自动调优问题。

Real-Time Auto-Optimization in Unknown Environments via Structure-Exploiting Dual Control for Exploration and Exploitation

论文配图:Real-Time Auto-Optimization in Unknown Environments via Structure-Exploiting Dual Control for Exploration and Exploitation
图 1 · 摘自论文原文
  • 利用奖励函数的凸-非线性结构,仅线性化非线性部分保留凸损失
  • 计算耗时降至83微秒,速度提升一个数量级
  • 适合嵌入式实时系统,如车载控制场景

本文针对未知环境中自动优化问题,提出一种快速数值双控方法(DCEE),以实现探索与利用的高效平衡。传统双控方法因依赖标准优化包或梯度更新,计算负担重。本文揭示了DCEE中奖励函数具有固有的凸-非线性结构:利用与探索项构成统一的非线性残差映射,外层为凸损失。据此,提出仅线性化非线性残差、保持凸外层的结构化求解方法,使每个子问题转化为可可靠求解的结构化凸形式。所得广义高斯-牛顿海森近似为半正定,仅依赖一阶导数,支持快速在线计算。在车辆巡航自动优化任务上评估,仿真与软硬件协同实验表明,该方法显著提升控制性能,计算时间最高仅83微秒,在典型车载嵌入式CPU上实现约十倍加速。

原文摘要 · Abstract (English)

This paper develops a fast numerical dual control for exploration and exploitation (DCEE) method to address auto-optimization problems in unknown environments. In auto-optimization problems, the optimal operating condition is unknown a priori and may vary with the environment. As in classical dual control techniques, computational burden remains a major concern in DCEE for active learning. Existing DCEE methods provide a principled exploration-exploitation objective, but mainly realized through standard optimization packages or explicit gradient-type update laws, where the numerical structure of the DCEE has not been fully exploited. This paper shows that the reward function in DCEE has an inherent convex-over-nonlinear structure, where the exploitation and exploration terms form a unified nonlinear residual map equipped with a convex outer loss. Benefiting from this structure, a structure-exploiting numerical method is developed by linearizing only the nonlinear residual map while preserving the convex outer loss. Thus, each subproblem is transformed into a structured convex form that can be solved reliably. The resulting generalized Gauss-Newton Hessian approximation is positive semidefinite and depends only on first-order derivatives, thereby supporting fast online computation. The proposed method is evaluated on a vehicle cruising auto-optimization problem and compared with existing methods. Simulation and hardware-in-the-loop experimental results show that the proposed method improves control performance and achieves a speedup of approximately one order of magnitude, with a microsecond-level maximum computation time of only 83 μs on a typical vehicle embedded CPU.

自动优化双控实时计算凸优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。