arXiv:2512.05107cs.RO2025-12被引 30

为视觉语言动作模型设计分阶段强化学习,提升机器人操作成功率

STARE-VLA: Progressive Stage-Aware Reinforcement for Fine-Tuning Vision-Language-Action Models

  • 将长序列动作分解为语义阶段,提供精准的阶段奖励信号
  • 在仿真环境中实现98.0%和96.4%的成功率,达当前最优
  • 适合研究机器人控制与强化学习融合的开发者参考

近期基于大语言模型和强化学习微调的视觉-语言-动作(VLA)模型在机器人操作任务中取得显著进展。现有方法常将长时序动作视为语言序列,采用轨迹级优化如轨迹偏好优化(TPO)或近端策略优化(PPO),导致信用分配粗糙且训练不稳定。然而,动作轨迹具有因果链式阶段结构,各阶段学习难度不同,不同于语言中语序可变而语义不变的特点。为此,本文提出阶段感知强化(STARE)模块,将长时序动作轨迹分解为语义有意义的阶段,并提供密集、可解释且阶段对齐的奖励信号。将STARE集成至TPO和PPO,分别得到阶段感知TPO(STA-TPO)和阶段感知PPO(STA-PPO),分别用于离线阶段偏好优化与在线阶段内交互。进一步以监督微调为初始化,提出模仿→偏好→交互(IPI)串联微调流程,提升动作准确性。在SimplerEnv与ManiSkill3上的实验表明,该方法取得98.0%和96.4%的最高成功率达,显著优于现有方法。

原文摘要 · Abstract (English)

Recent advances in Vision-Language-Action (VLA) models, powered by large language models and reinforcement learning-based fine-tuning, have shown remarkable progress in robotic manipulation. Existing methods often treat long-horizon actions as linguistic sequences and apply trajectory-level optimization methods such as Trajectory-wise Preference Optimization (TPO) or Proximal Policy Optimization (PPO), leading to coarse credit assignment and unstable training. However, unlike language, where a unified semantic meaning is preserved despite flexible sentence order, action trajectories progress through causally chained stages with different learning difficulties. This motivates progressive stage optimization. Thereby, we present Stage-Aware Reinforcement (STARE), a module that decomposes a long-horizon action trajectory into semantically meaningful stages and provides dense, interpretable, and stage-aligned reinforcement signals. Integrating STARE into TPO and PPO, we yield Stage-Aware TPO (STA-TPO) and Stage-Aware PPO (STA-PPO) for offline stage-wise preference and online intra-stage interaction, respectively. Further building on supervised fine-tuning as initialization, we propose the Imitation -> Preference -> Interaction (IPI), a serial fine-tuning pipeline for improving action accuracy in VLA models. Experiments on SimplerEnv and ManiSkill3 demonstrate substantial gains, achieving state-of-the-art success rates of 98.0 percent on SimplerEnv and 96.4 percent on ManiSkill3 tasks.

机器人控制强化学习多模态动作规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。