用离线演示生成探索令牌,提升视觉语言动作模型的采样效率
ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning

- 基于离线演示生成离散行为先验,引导策略探索多样化行为模式
- 在有限交互次数下显著提升状态-动作覆盖范围和任务完成率
- 适合机器人操控等需高效探索的现实场景,尤其适用于数据受限环境
强化学习在复杂操作任务中展现巨大潜力,但其实际可扩展性受制于环境交互成本。本文发现当前视觉-语言-动作强化学习框架存在探索停滞瓶颈,轨迹多样性比样本数量更关键。为此提出ExToken框架,将离线演示生成的离散行为先验作为条件,引导策略进行结构化探索,显著提升状态-动作覆盖与探索效率。为衔接训练探索与部署时确定性推理,引入状态相关令牌选择器,自适应预测未见场景的有效行为模式。在仿真与真实机器人操作任务中的大量实验表明,ExToken持续加速收敛、提升性能,并在极低交互预算下表现出强鲁棒性。
原文摘要 · Abstract (English)
Reinforcement Learning (RL) has demonstrated significant potential for improving Vision-Language-Action (VLA) models on complex manipulation tasks. However, its practical scalability remains severely limited by the substantial cost of environmental interactions. In this work, we first investigate the exploration stagnation bottleneck in current VLA-RL frameworks and reveal that trajectory diversity is fundamentally more important to sample efficiency than the sheer quantity of collected rollouts. Motivated by these insights, we introduce RL Exploration Token (ExToken), a simple yet general framework that condition VLA policies on discrete behavioral priors derived from offline demonstrations for structured exploration. By conditioning the policy on different tokens during rollout collection, ExToken encourages the agent to explore diverse behavioral modes, substantially improving state-action coverage and exploration efficiency. To bridge exploration during training with deterministic inference at deployment, ExToken further incorporates a state-conditioned token selector that adaptively predicts effective behavioral modes for unseen scenarios. Extensive experiments across simulated and real-world robotic manipulation tasks demonstrate that ExToken consistently accelerates convergence, improves task performance, and exhibits strong robustness under highly constrained interaction budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。