arXiv:2511.01331cs.ROcs.LG2025-11被引 14

提升视觉语言动作模型在扰动下的可靠性,通过强化学习显式增强鲁棒性。

RobustVLA: Robustness-Aware Reinforcement Post-Training for Vision-Language-Action Models

  • 引入雅可比与平滑正则化,降低对观测噪声和执行扰动的敏感度。
  • 在多种机器人环境中,鲁棒性显著优于现有最先进方法。
  • 适合关注实际部署中模型可靠性的机器人研究者与工程师。

视觉语言动作(VLA)模型近年来成为机器人操作的强大通用策略,得益于大规模多模态预训练。然而,在分布外部署中,不可避免的干扰如观测噪声、传感器误差或执行扰动普遍存在时,其泛化能力往往不可靠。尽管基于强化学习(RL)的后训练提供了一种实用的适应方式,但现有方法主要聚焦于奖励最大化,忽视了对环境不确定性的鲁棒性。本文提出RobustVLA,一种轻量级在线强化学习后训练方法,旨在显式增强VLA模型的抗扰能力。通过系统性鲁棒性分析,我们识别出两项关键正则化:雅可比正则化(缓解观测噪声敏感性)与平滑性正则化(稳定动作扰动下的策略)。在多样机器人环境中的大量实验表明,RobustVLA在鲁棒性和可靠性方面显著优于先前最先进方法。结果强调了有原则的鲁棒性感知强化学习后训练对于提升VLA模型可靠性和鲁棒性的关键作用。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have recently emerged as powerful general-purpose policies for robotic manipulation, benefiting from large-scale multi-modal pre-training. However, they often fail to generalize reliably in out-of-distribution deployments, where unavoidable disturbances such as observation noise, sensor errors, or actuation perturbations become prevalent. While recent Reinforcement Learning (RL)-based post-training provides a practical means to adapt pre-trained VLA models, existing methods mainly emphasize reward maximization and overlook robustness to environmental uncertainty. In this work, we introduce RobustVLA, a lightweight online RL post-training method designed to explicitly enhance the resilience of VLA models. Through a systematic robustness analysis, we identify two key regularizations: Jacobian regularization, which mitigates sensitivity to observation noise, and smoothness regularization, which stabilizes policies under action perturbations. Extensive experiments across diverse robotic environments demonstrate that RobustVLA significantly outperforms prior state-of-the-art methods in robustness and reliability. Our results highlight the importance of principled robustness-aware RL post-training as a key step toward improving the reliability and robustness of VLA models.

机器人强化学习鲁棒性多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。