用强化学习优化视觉动作分块模型,提升工业抓取的稳定性和安全性。
PAC-ACT: Post-training Actor-Critic for Action Chunking Transformers

- 在动作分块级别进行强化学习优化,保持实时控制性能。
- 接触力峰值降低,超过60N的力读数减少46倍。
- 适合需要低延迟、高安全性的工业机器人场景。
精密工业接触操作需在位姿扰动和接触力约束下保持可靠机器人策略。视觉-语言-动作模型虽具强泛化能力,但推理延迟高、显存消耗大;而视觉-动作分块策略更适配实时工业控制,却常因行为克隆训练导致接触密集任务中的分布偏移。本文提出PAC-ACT,一种针对预训练动作分块变压器(Action Chunking Transformer)的强化学习后训练框架。PAC-ACT在分块层级重构策略优化,构建基于ACT转移的演员-评论家架构,并引入混合行为先验约束,在在线微调中保留预训练动作分布。在工业精密接触基准测试中,PAC-ACT显著提升任务成功率、接触稳定性与力安全性能,同时保持低延迟与低显存占用。在轮廓任务中,峰值接触力下降,超过60N的力读数比例降低46倍。稀疏奖励消融实验表明,所提行为先验约束可有效支持随机初始位姿下的探索。
原文摘要 · Abstract (English)
Precision industrial contact manipulation requires reliable robot policies under pose perturbations and contact-force constraints. Vision-language-action models offer broad generalization but often introduce high inference latency and GPU-memory cost, while vision-action chunking policies are more suitable for real-time industrial control. However, these policies are usually trained by behavior cloning and suffer from distribution shift in contact-rich tasks. This paper proposes PAC-ACT, a reinforcement-learning post-training framework for pretrained Action Chunking Transformer policies. PAC-ACT reformulates policy optimization at the chunk level, constructs an ACT-transferred actor-critic architecture, and introduces a hybrid behavior-prior constraint to preserve the pretrained action distribution during online fine-tuning. Experiments on industrial precision-contact benchmarks show that PAC-ACT improves task success, contact stability, and force safety while retaining low latency and low GPU-memory usage. On the Contour task, PAC-ACT significantly reduces peak contact force and decreases the proportion of force readings above 60 N by 46 times. Sparse-reward ablations further show that the proposed behavior-prior constraint enables effective exploration under randomized initial poses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。