arXiv:2606.26025cs.ROcs.CV2026-06

让机器人通过自动生成的交互数据,自己理解当前环境和自身状态。

In-Context World Modeling for Robotic Control

论文配图:In-Context World Modeling for Robotic Control
图 1 · 摘自论文原文
  • 用短期自生成交互数据推断系统变量,实现无需参数更新的适应。
  • 在新摄像头视角下,性能显著优于传统视觉语言动作模型。
  • 适合需要快速适应新环境的机器人控制场景。

现代视觉-语言-动作(VLA)模型在面对新设置(如摄像头视角变化或机器人形态不同)时难以泛化,因其通常仅依赖当前观测和语言指令,忽略了系统配置这一变量。此类模型隐含假设训练时遇到的执行上下文是固定的,导致在新环境中需大量微调。本文提出一种称为上下文世界建模(ICWM)的框架,将系统识别视为上下文自适应问题。ICWM使机器人策略能从少量自生成的、任务无关的交互历史中自主推断关键系统变量。不同于传统上下文学习使用示范来指定任务,ICWM利用上下文窗口理解系统运行机制。通过在任务执行前处理这些交互,模型隐式捕捉当前系统的动态特性,从而在不进行参数更新的情况下适配新配置。仿真与真实机器人平台的大量实验表明,ICWM在新摄像头视角下的表现显著优于标准VLA基线。

原文摘要 · Abstract (English)

Modern Vision-Language-Action (VLA) models often fail to generalize to novel setups, such as altered camera viewpoints or robot morphologies, because they are typically conditioned only on current observations and language instructions. By ignoring the underlying system configuration as a variable, these models implicitly assume a fixed execution context encountered during training, necessitating data-intensive fine-tuning for any new environment. In this work, we introduce In-Context World Modeling (ICWM), a framework that treats system identification as an in-context adaptation problem. ICWM enables robot policies to autonomously infer essential system variables from a short history of self-generated, task-agnostic interactions. Unlike traditional In-Context Learning that uses demonstrations to specify what task to perform, ICWM leverages the context window to understand how the system operates. By processing these interactions before task execution, the model implicitly captures the world dynamics of the current system, enabling adaptation to novel configurations without parameter updates. Extensive experiments in simulation and on real-world robot platforms demonstrate that ICWM significantly outperforms standard VLA baselines on novel camera viewpoints.

机器人控制上下文学习世界建模自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。