arXiv:2602.17259cs.RO2026-02被引 10

通过多未来表征对齐,提升机器人政策的世界感知能力。

FRAPPE: Infusing World Modeling into Generalist Policies via Multiple Future Representation Alignment

  • 分两阶段微调:先预测未来隐变量,再并行对齐多视觉模型表征
  • 在RoboTwin和真实任务中实现长时程与未知场景的强泛化性能
  • 减少对动作标注数据依赖,提升微调效率,适合通用机器人策略

让视觉-语言-动作模型具备环境动态预测能力(即世界建模),对提升机器人推理与泛化至关重要。但现有方法存在两大问题:一是训练目标过度强调像素级重建,限制语义学习;二是推理时依赖预测未来观测,易导致误差累积。为此,本文提出基于并行渐进扩展的未来表征对齐方法(FRAPPE)。该方法采用两阶段微调:中期训练中模型学习预测未来观测的隐表示;后期则并行扩展计算负载,同时与多个不同的视觉基础模型对齐表征。该方法显著提升微调效率,降低对动作标注数据的依赖,为通用机器人策略增强世界感知提供了可扩展、数据高效的新路径。在RoboTwin基准和真实任务上的实验表明,FRAPPE优于当前最优方法,在长时程及未见场景中展现强大泛化能力。

原文摘要 · Abstract (English)

Enabling VLA models to predict environmental dynamics, known as world modeling, has been recognized as essential for improving robotic reasoning and generalization. However, current approaches face two main issues: 1. The training objective forces models to over-emphasize pixel-level reconstruction, which constrains semantic learning and generalization 2. Reliance on predicted future observations during inference often leads to error accumulation. To address these challenges, we introduce Future Representation Alignment via Parallel Progressive Expansion (FRAPPE). Our method adopts a two-stage fine-tuning strategy: In the mid-training phase, the model learns to predict the latent representations of future observations; In the post-training phase, we expand the computational workload in parallel and align the representation simultaneously with multiple different visual foundation models. By significantly improving fine-tuning efficiency and reducing dependence on action-annotated data, FRAPPE provides a scalable and data-efficient pathway to enhance world-awareness in generalist robotic policies. Experiments on the RoboTwin benchmark and real-world tasks demonstrate that FRAPPE outperforms state-of-the-art approaches and shows strong generalization in long-horizon and unseen scenarios.

世界建模机器人表征对齐泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。