arXiv:2512.03538cs.RO2025-12被引 2

让通用世界模型变专精,提升机器人操控成功率

AdaPower: Specializing World Foundation Models for Predictive Manipulation

  • 推理时动态调整模型,提升对任务的适应性
  • 在LIBERO基准上任务成功率提升超41%,无需重训策略
  • 轻量改造适合部署,兼顾效率与泛化能力

世界基础模型(WFMs)具备出色的视觉动态模拟能力,但在精确机器人控制中的应用受限于生成真实感与控制精度之间的差距。现有方法将WFMs用作合成数据生成器,存在计算成本高、预训练视觉语言策略利用不足的问题。本文提出AdaPower(Adaapt and Empower),一种轻量级适配框架,通过两个新组件:时间-空间测试时训练(TS-TTT)实现推理时模型自适应,以及记忆持久性(MP)保障长时序一致性。集成于模型预测控制框架中,适配后的世界模型显著提升了预训练视觉语言策略(VLA)的性能,在LIBERO基准上任务成功率提升超过41%,且无需策略重训练,同时保持计算高效与通用能力。

原文摘要 · Abstract (English)

World Foundation Models (WFMs) offer remarkable visual dynamics simulation capabilities, yet their application to precise robotic control remains limited by the gap between generative realism and control-oriented precision. While existing approaches use WFMs as synthetic data generators, they suffer from high computational costs and underutilization of pre-trained VLA policies. We introduce \textbf{AdaPower} (\textbf{Ada}pt and Em\textbf{power}), a lightweight adaptation framework that transforms general-purpose WFMs into specialist world models through two novel components: Temporal-Spatial Test-Time Training (TS-TTT) for inference-time adaptation and Memory Persistence (MP) for long-horizon consistency. Integrated within a Model Predictive Control framework, our adapted world model empowers pre-trained VLAs, achieving over 41\% improvement in task success rates on LIBERO benchmarks without policy retraining, while preserving computational efficiency and generalist capabilities.

机器人控制世界模型轻量适配MPC

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。