不修改模型,用动态反馈修正动作指令,提升机械臂操作精度。
APEX: Adaptive Policy Execution for Precise Manipulation

- 在策略与控制器间插入自适应模块,实时重构可行动作参考
- 测试时根据底层状态反馈调整,使执行误差降低41.2%
- 无需改动原模型,适用于各类视觉-动作策略,尤其适合高精度任务
当前模仿学习方法(如视觉-动作与视觉-语言-动作策略)通常输出高层动作指令,由底层控制器执行。由于缺乏高阶参考信号,且策略训练时未感知底层控制动态,导致实际执行动作与指令存在系统性偏差,严重影响高精度操作。现有方法需修改策略架构或底层控制器,均需侵入式更改预训练模型。本文提出APEX:一种可即插即用的自适应策略执行框架,插入于策略与控制器之间,基于策略输出和低层状态反馈动态重构可行参考,并在测试时自适应调整,具备收敛性保证。大量实验证明,该方法在演示回放中将控制器引起的跟踪误差降低41.2%,并在四类视觉-动作与视觉-语言-动作策略上,使操作成功率提升4.8%至25.8个百分点。
原文摘要 · Abstract (English)
Modern imitation learning methods, including visuomotor and Vision-Language-Action (VLA) policies, typically output high-level action references that are executed by low-level controllers. However, the absence of higher-order reference signals, together with the policy's lack of awareness of the underlying low-level control dynamics during training, inevitably induces an execution gap. As a result, realized actions deviate systematically from policy-commanded ones, with a critical impact on precision-sensitive manipulation. Prior work either modifies the policy architecture or the low-level controller, both requiring intrusive changes to the pretrained policy or packaged controller. This raises a natural question: when the policy and controller are both treated as inaccessible black boxes, can we bridge the execution gap? We propose Adaptive Policy Execution (APEX), a plug-and-play framework inserted between the policy and the controller that reconstructs a dynamically feasible reference from policy outputs and adapts at test-time according to low-level state feedback, with a provable convergence guarantee. Extensive empirical studies show that APEX reduces controller-induced tracking error by 41.2% on demonstration replay and improves manipulation success by 4.8--25.8 percentage points across four visuomotor and VLA policy classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。