提出新方法提升流模型在线强化学习性能,解决探索空间大时的精度问题。
$π$-StepNFT: Wider Space Needs Finer Steps in Online RL for Flow-based VLAs
- 仅需单次前向传播完成优化,无需价值网络和似然计算
- 在LIBERO上实现强少样本鲁棒性,在ManiSkill上优于基线模型
- 适合复杂现实场景,尤其对分布外泛化有显著提升
基于流的视觉-语言-动作(VLA)模型在具身控制中表现优异,但在多步采样时面临不可处理的似然计算问题,阻碍了在线强化学习。我们提出π-StepNFT(逐步负感知微调),一种无需评判器和似然计算的框架,每轮优化仅需一次前向传播。我们发现更宽的探索空间需要更精细的逐步引导以实现对齐。实验表明,π-StepNFT在LIBERO上解锁了潜在性能,具备竞争力的少样本鲁棒性;在ManiSkill上实现更优泛化能力,优于基于价值的基线模型,在分布外(OOD)场景下有效防止对多模态特征的过拟合。该方法为复杂真实应用提供了可扩展解决方案。
原文摘要 · Abstract (English)
Flow-based vision-language-action (VLA) models excel in embodied control but suffer from intractable likelihoods during multi-step sampling, hindering online reinforcement learning. We propose \textbf{\textit{$\boldsymbolπ$-StepNFT}} (Step-wise Negative-aware Fine-Tuning), a critic-and-likelihood-free framework that requires only a single forward pass per optimization step and eliminates auxiliary value networks. We identify that wider exploration spaces necessitate finer-grained, step-wise guidance for alignment. Empirically, $π$-StepNFT unlocks latent potential on LIBERO with competitive few-shot robustness. Moreover, it achieves superior generalization on ManiSkill, outperforming value-based baselines in OOD scenarios by preventing overfitting to multimodal features. This property offers a scalable solution promising for complex real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。