arXiv:2605.06877cs.LG2026-05

用自注意力机制动态调节机械臂控制增益,解决摩擦记忆不可观测难题。

Temporal Attention for Adaptive Control of Euler-Lagrange Systems with Unobservable Memory

论文配图:Temporal Attention for Adaptive Control of Euler-Lagrange Systems with Unobservable Memory
图 1 · 摘自论文原文
  • 通过短期运动历史的自注意力块生成控制器增益,应对不可观测的摩擦记忆。
  • 在短时和匹配记忆下,误差降低12%和19%,效应量达-1.1和-2.1,显著优于基线。
  • 适合需要高精度轨迹跟踪的机器人控制场景,尤其关注非线性摩擦系统。

当欧拉-拉格朗日系统中的摩擦由有限时域内部状态决定且无法从关节测量中直接获取时,自适应控制面临挑战。此时闭环状态不再满足马尔可夫性,标准确定性等价自适应律可能失去收敛性保证。本文提出一种元控制架构:计算力矩控制器的增益由处理近期运动历史的一层自注意力模块生成。注意力头数量在策略训练前通过梯度自协方差的代理分析预先选定,该分析基于作者此前提出的增量秩追踪框架的时序扩展。选定头数作为固定超参数用于强化学习阶段,在屏蔽可接受性约束下训练策略。实验在带非线性摩擦与变负载的2自由度机械臂上进行。在短时与匹配记忆情形下,单层注意力元控制器性能优于更深的Transformer基线,分别实现12%与19%的跟踪误差下降,效应量约为-1.1和-2.1,且曼惠特尼检验p < 0.05。但在长记忆情形下优势消失,四次训练中出现发散或负载无关策略崩溃,暴露出静态预设头数的缺陷。因此提出将秩追踪引入强化学习循环,允许运行时动态增删注意力头。

原文摘要 · Abstract (English)

Adaptive control of Euler-Lagrange systems is challenging when friction is governed by a finite-horizon internal state that is not directly observable from joint measurements. In this setting, the measured closed-loop state is no longer Markovian, and standard certainty-equivalence adaptive laws may lose their convergence guarantees. The paper proposes a meta-control architecture in which the gains of a computed-torque controller are generated by a self-attention block processing a short window of recent motion history. The number of attention heads is selected before policy training through a surrogate analysis of the autocovariance of the memory-state gradient along the temporal window. This surrogate is based on a temporal adaptation of an incremental rank-tracking framework previously developed by the authors. The selected head count is then fixed and used as an architectural hyperparameter in a reinforcement-learning stage, where the policy is trained under a shielded admissibility constraint. The approach is tested on a 2-DOF manipulator with nonlinear friction and variable payload. In the short and matched memory regimes, the single-layer attention-only meta-controller outperforms a deeper Transformer baseline, with tracking-error reductions of 12 and 19 percentage points, respectively. The reported effect sizes are large, with d approximately -1.1 and -2.1, and Mann-Whitney p < 0.05 in both cases. In the long memory regime, however, the advantage disappears. Four out of ten training runs show either divergence or payload-invariant policy collapse, revealing a weakness in the static Phase-1 head-count prescription. This motivates moving rank-tracking inside the reinforcement-learning loop, allowing attention heads to be pruned or grown at runtime instead of fixed before training.

自适应控制注意力机制机器人动力学强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。