arXiv:2607.26370cs.ROcs.LG2026-07

自适应学习控制可无悔追踪未知动态,支持结构化、随机与对抗性运动切换。

Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret

论文配图:Self-Adaptive Learning and Model Predictive Control for Tracking Unknown Dynamics with No Regret
图 1 · 摘自论文原文
  • 从零开始自监督学习多个预测器,动态选择最优匹配目标行为
  • 在无误差无切换时实现渐近最优,有误差或切换时平均遗憾与学习误差成正比
  • 适用于动态地图、交通管控等需追踪或避障的机器人场景

我们提出一种用于追踪未知目标动态的自适应在线学习控制方法。目标动态可能呈现切换行为,包括结构化、随机和/或对抗性运动。此类挑战性追踪场景出现在动态地图、交通控制和追逃任务中,机器人需追踪、追逐或避免与动态地标、物体、人类等发生碰撞,其运动规律未知。本方法通过自监督、一次性、计算高效的学习方式,从零开始同时学习多个预测器,并自适应选择最匹配观测行为的预测器。该方法在期望下具备有限时间近似最优性保证,其性能由目标动态学习误差和切换频率决定。在无误差且无切换情况下,方法渐近匹配已知目标动态的非因果最优控制策略,即期望无悔;在存在学习误差和切换时,性能可平稳退化——如仅有误差无切换,平均遗憾与平均学习误差及切换次数成正比。为证明这些保证,需采用区别于现有基于RFF的在线学习方法的新技术路径。我们在Crazyflie仿真与硬件实验中验证了该方法,在从结构化到随机再到对抗性目标轨迹的多场景下,对比了非随机、核方法和神经网络在线学习方法,表现优异。

原文摘要 · Abstract (English)

We propose a self-adaptive online learning for control method for tracking unknown target dynamics. The target dynamics can exhibit switching behavior, particularly, a mixture of structured, random, and/or adversarial motion. Such challenging target tracking scenarios arise in applications of dynamic mapping, traffic control, and pursuit evasion, where robots need to track, pursue, or avoid collision with moving landmarks, objects, humans, etc., whose dynamics are unknown. Our method simultaneously learns multiple predictors from scratch, via self-supervised, one-shot, and computationally efficient learning, and adaptively selects the best one to match the observed target behavior. The method enjoys finite-time near-optimality guarantees in expectation, characterized as a function of the learning error of the target dynamics and the frequency that the target dynamics switch. In the absence of both error and switching, the method asymptotically matches the optimal non-causal control policy that knows a priori the target dynamics, i.e., the method enjoys no regret in expectation. In the presence of learning errors and switching, the method degrades gracefully, \eg when there are errors and no switching, the average regret is proportional to the average learning error and switching times. To prove these guarantees, a novel technical approach is required compared to the existing works that employ RFF-based online learning. We validate our method in Crazyflie simulations and hardware experiments, across target trajectories that vary from structured to random to adversarial, in comparison to non-stochastic, kernel-based, and neural-network-based methods for online learning.

自适应控制在线学习机器人追踪无悔学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。