用世界模型和动作奖励微调,提升视觉语言动作模型的可靠性与泛化能力。
NORA-1.5: A Vision-Language-Action Model Trained using World Model- and Action-based Preference Rewards
- 基于流匹配构建动作专家,结合世界模型评估动作是否导向目标。
- 在仿真与真实机器人上均超越多个主流VLA模型,任务成功率显著提升。
- 适合希望提升智能体可靠性的研究者或工业部署团队使用。
视觉-语言-动作(VLA)模型在多种具身任务中表现优异,但在跨平台部署和真实环境中的可靠性与泛化能力仍不足。本文提出NORA-1.5,基于预训练的NORA骨干网络,引入基于流匹配的动作专家。该架构改进使NORA-1.5在仿真与真实世界基准测试中均优于NORA及多个先进VLA模型。为增强鲁棒性与任务成功率,我们设计了一组后训练奖励模型:(i) 动作条件的世界模型(WM),评估生成动作是否导向目标;(ii) 偏离真实轨迹的启发式判别器,区分优劣动作。利用这些奖励信号构建偏好数据集,并通过直接偏好优化(DPO)适配目标具身形态。大量实验表明,基于奖励的后训练在仿真与真实机器人环境中均持续提升性能,证明简单而有效的奖励模型可显著提升VLA模型的可靠性。本研究强调了NORA-1.5与奖励引导后训练在真实世界部署中的可行性。
原文摘要 · Abstract (English)
Vision--language--action (VLA) models have recently shown promising performance on a variety of embodied tasks, yet they still fall short in reliability and generalization, especially when deployed across different embodiments or real-world environments. In this work, we introduce NORA-1.5, a VLA model built from the pre-trained NORA backbone by adding to it a flow-matching-based action expert. This architectural enhancement alone yields substantial performance gains, enabling NORA-1.5 to outperform NORA and several state-of-the-art VLA models across both simulated and real-world benchmarks. To further improve robustness and task success, we develop a set of reward models for post-training VLA policies. Our rewards combine (i) an action-conditioned world model (WM) that evaluates whether generated actions lead toward the desired goal, and (ii) a deviation-from-ground-truth heuristic that distinguishes good actions from poor ones. Using these reward signals, we construct preference datasets and adapt NORA-1.5 to target embodiments through direct preference optimization (DPO). Extensive evaluations show that reward-driven post-training consistently improves performance in both simulation and real-robot settings, demonstrating significant VLA model-reliability gains through simple yet effective reward models. Our findings highlight NORA-1.5 and reward-guided post-training as a viable path toward more dependable embodied agents suitable for real-world deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。