arXiv:2603.10422cs.CV2026-03被引 4

用隐空间动态替代像素监督,让机器人模型更稳定地学习动作。

World2Act: Latent Action Post-Training from World Model Dynamics

  • 通过对比对齐视频与动作隐空间,构建共享表示
  • 在仿真和真实机器人上分别提升2.5%和6.7%成功率
  • 避免像素级误差干扰,适合追求鲁棒性的视觉-语言-动作模型

世界模型(WM)为视觉-语言-动作(VLA)策略的后训练提供了动态先验,有助于提升任务与场景变化下的泛化能力。然而,多数基于WM的后训练方法依赖像素空间监督,使策略对不完美模拟生成的视觉伪影敏感。本文提出World2Act,一种无需像素空间监督的隐空间后训练框架,将WM动态迁移至VLA策略。该方法分两阶段:1)通过对比对齐WM动态隐状态与动作嵌入,构建共享视频-动作隐空间;2)引导策略的动作表示向WM模拟的动态靠拢,而非解码后的像素。基于GR00T-N1.6,World2Act在仿真基准(RoboCasa、LIBERO、Bridge-SIMPLER)上实现最高+2.5%的绝对成功率提升,在真实机器人上提升+6.7%,显著优于微调基线。尤其在LIBERO上,其性能比像素空间监督高出+6.0%,且后者反而降低基线表现,表明隐空间动态是更稳定的后训练方式。

原文摘要 · Abstract (English)

World Models (WMs) offer a promising mechanism for post-training Vision-Language-Action (VLA) policies by providing dynamics priors that improve generalization under task and scene variation. However, most WM-based post-training methods rely on pixel-space supervision, making policies sensitive to visual artifacts introduced by imperfect WM rollouts. We present World2Act, a latent-space post-training framework that transfers WM dynamics to the VLA policy without pixel-space supervision. World2Act operates in two stages: 1) it induces a shared video-action latent space by contrastively aligning WM-dynamics latents with action embeddings, and 2) it post-trains the VLA by guiding policy action representations toward WM-imagined dynamics rather than decoded pixels. Built on GR00T-N1.6, World2Act delivers absolute success-rate gains of up to +2.5% on simulation benchmarks (RoboCasa, LIBERO, Bridge-SIMPLER) and +6.7% on a real robot over finetuned VLA baselines. Notably, it outperforms pixel-space WM supervision by up to +6.0%, including on LIBERO where pixel supervision degrades the baseline, suggesting that latent WM dynamics offer a more stable WM-based post-training alternative to pixel-space transfer.

世界模型后训练动作学习隐空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。