让视觉模型在想象中更可靠,通过动态调整风格和隐空间保持稳定性。
Sword: Style-Robust World Models as Simulators via Dynamic Latent Bootstrapping for VLA Policy Post-Training

- 分离视觉风格与任务动态,提升环境泛化能力
- 动态隐空间重初始化,减少长程误差累积
- 适合需要高鲁棒性的视觉-语言-动作策略后训练
将视觉-语言-动作(VLA)模型与世界模型结合日益受到关注。一种代表性方法将学习到的世界模型作为生成模拟器,实现完全在“想象”中进行策略优化。然而,在如LIBERO基准等特定环境中部署时,现有世界模型常面临泛化能力差和长程误差累积问题。封闭环回放过程中,模型对初始状态扰动高度敏感;颜色、光照等微小变化可引发级联幻觉,导致严重模糊或过曝。此外,长程误差累积进一步降低预测未来状态的质量与保真度。为此,我们提出Sword框架。该方法引入结构引导的风格增强,解耦交互环境的视觉纹理与任务相关动态,从而提升泛化性。同时提出动态隐空间重初始化机制,在保持训练与推理一致性的同时控制内存开销。在LIBERO基准上的大量实验表明,相比基线方法WoVR,Sword在泛化性、生成质量、鲁棒性、保真度及强化学习后训练的成功率方面均有显著提升。
原文摘要 · Abstract (English)
The integration of Vision-Language-Action (VLA) models with World Models has gained increasing attention. One representative approach treats learned World Models as generative simulators, enabling policy optimization entirely within "imagination." However, when deployed as simulators for specific environments such as the LIBERO benchmark, existing World Models often suffer from poor generalization and long-horizon error accumulation. During closed-loop rollouts, these models are highly sensitive to initial-state perturbations; minor changes in color, illumination, and other visual factors can trigger cascading hallucinations, leading to severe blurriness or overexposure. Moreover, long-horizon error accumulation further degrades the quality and fidelity of predicted future states. These issues limit the reliability of World Models as simulators. To mitigate these problems, we propose Sword, a robust World Model framework. Our method introduces Structure-Guided Style Augmentation to disentangle the visual textures of interactive environments from task-relevant dynamics, thereby improving generalization. We further propose Dynamic Latent Bootstrapping, which maintains consistency between training and inference while keeping memory consumption low. Extensive experiments on the LIBERO benchmark show that our method significantly outperforms the baseline WoVR in terms of generalization, generation quality, robustness, fidelity, and the success rate of reinforcement-learning post-training for VLA models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。