让机器人动作更容错,提升训练效率和泛化能力
Boosting Vision-Language-Action Finetuning with Feasible Action Neighborhood Prior
- 引入可行动作邻域先验,使模型输出更平滑一致
- 在多种场景下样本效率提升,成功率达90%以上
- 适合需要高效、鲁棒动作训练的机器人研究者
在真实世界机器人操作中,状态通常对应一个近似等效的动作邻域(FAN),即多个动作可产生无差别的进展。然而,主流视觉-语言-动作(VLA)训练方法沿用语言模型范式,未利用这一特性,导致泛化性差、样本效率低。本文提出一种基于FAN的正则化方法,通过高斯先验引导模型输出在优选方向与幅度附近保持局部平滑且单峰,从而匹配物理操作的内在容错性。在强化微调(RFT)和监督微调(SFT)的大量实验中,该方法显著提升了样本效率,在分布内与分布外(OOD)场景下均实现超过90%的成功率。此正则化策略为高效、可泛化的VLA适应提供了原理清晰且实用的解决方案。
原文摘要 · Abstract (English)
In real-world robotic manipulation, states typically admit a neighborhood of near-equivalent actions. That is, for each state, there exist a feasible action neighborhood (FAN) rather than a single correct action, within which motions yield indistinguishable progress. However, prevalent VLA training methodologies are directly inherited from linguistic settings and do not exploit the FAN property, thus leading to poor generalization and low sample efficiency. To address this limitation, we introduce a FAN-guided regularizer that shapes the model's output distribution to align with the geometry of FAN. Concretely, we introduce a Gaussian prior that promotes locally smooth and unimodal predictions around the preferred direction and magnitude. In extensive experiments across both reinforced finetuning (RFT) and supervised finetuning (SFT), our method achieves significant improvement in sample efficiency, and success rate in both in-distribution and out-of-distribution (OOD) scenarios. By aligning with the intrinsic action tolerance of physical manipulation, FAN-guided regularization provides a principled and practical method for sample-efficient, and generalizable VLA adaptation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。