arXiv:2604.14484cs.ROcs.AI2026-04

揭示控制器增益如何放大行为克隆误差,影响机器人控制可靠性。

Behavior Cloning Under PD Control: A Finite-Horizon Theory of Gain-Dependent Error Amplification

  • 分析了PD控制器增益对行为克隆误差的有限时域放大机制。
  • 增益越强(刚度高)误差方差越大,阻尼越大则误差越小。
  • 适用于评估机器人控制中增益选择的优劣,尤其对高精度任务重要。

在位置控制机器人上进行行为克隆时,策略动作受PD控制回路执行。本文给出了一个有限时域、非渐近的分析,揭示控制器增益如何影响行为克隆失败。独立的子高斯动作误差通过增益相关的闭环动态传播为子高斯位置误差,最终导致的失败尾部由控制器放大倍数乘以验证损失和泛化松弛共同决定,因此仅看验证损失会误判增益优劣。在保持形状的上界假设下,分析分离出标签难度、注入强度与收缩性,表明柔顺-过阻尼增益最紧致,刚硬-欠阻尼最松散,混合模式则依赖系统特性。在标准标量二阶PD系统中,稳态位置误差方差随刚度增加而上升,随阻尼增加而下降,在稳定范围内成立;精确零阶保持离散化在主导阶次上继承该排序。该结果将Bronars等(2026)的误差衰减解释扩展至有限时域失败边界。

原文摘要 · Abstract (English)

Behavior cloning (BC) on position-controlled robots is shaped by the PD loop that executes policy actions. We give a finite-horizon, nonasymptotic analysis of how controller gains affect BC failure. Independent sub-Gaussian action errors propagate through gain-dependent closed-loop dynamics into sub-Gaussian position errors. The resulting failure tail is controlled by controller amplification multiplied by validation loss and generalization slack, so validation loss alone can mis-rank gains. Under shape-preserving upper-bound assumptions, the analysis separates label difficulty, injection strength, and contraction, ranking compliant-overdamped gains as tightest and stiff-underdamped gains as loosest, with the mixed regimes system-dependent. In the canonical scalar second-order PD system, stationary position-error variance increases with stiffness and decreases with damping over the stable range, and exact zero-order-hold discretization inherits the ordering to leading order. This extends the error-attenuation explanation of bronars et al. (2026) to finite-horizon failure bounds.

行为克隆控制理论机器人误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。