arXiv:2608.05799cs.ROcs.CV2026-08

提出新测试框架,揭示世界模型难以泛化到未见机器人形态。

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?

论文配图:XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments?
图 1 · 摘自论文原文
  • 构建跨形态测试环境,隔离机器人本体差异
  • 模型主要依赖视觉相似性而非物理相似性,泛化能力差
  • 需像素级动作和时空对齐才能零样本渲染新机器人

行动条件世界模型是机器人操作的有前途的可学习模拟器,但仅在训练过的机器人上评估无法揭示其是否捕捉物理动态或仅记忆视觉模式。为回答模型能否忠实呈现从未见过的机器人,我们引入XEWorld,一个控制性的跨形态测试平台,通过在物理相同的场景中评估保留的机器人来隔离本体差异。系统性分析发现共享架构瓶颈:当前模型主要作为2D视觉模式匹配器,其泛化由视觉相似性决定而非物理运动学相似性。受此限制,它们难以将抽象的数值关节动作转化为连贯的视觉轨迹,并无法从静态初始观测预测动态视觉变化。因此,成功零样本渲染未见形态严格依赖高度接地的线索,特别是像素空间动作和显式时空对齐。即使通过少样本适应绕过零样本障碍,强制外观恢复也会引发对已见形态的灾难性遗忘。这些失败暴露了将学习到的物理动态应用于新视觉外观的关键能力缺失,表明实现真正的跨形态泛化需要解耦视觉外观与底层物理动态的架构创新。

原文摘要 · Abstract (English)

Action-conditioned world models are promising learned simulators for robotic manipulation, yet evaluating them exclusively on training robots fails to reveal whether they capture physical dynamics or merely memorize visual patterns. To answer whether a model can faithfully render a robot it has never seen, we introduce XEWorld, a controlled cross-embodiment testbed for world models that isolates embodiments by evaluating held-out robots within physically identical scenes. Our systematic analysis uncovers a shared architectural bottleneck: current models act primarily as 2D visual pattern matchers whose generalization is governed by visual similarity rather than physical kinematic similarity. Driven by this limitation, they struggle to translate abstract numeric joint actions into coherent visual trajectories, and fail to predict dynamic visual changes from static initial observations. Consequently, successfully rendering an unseen embodiment zero-shot strictly requires heavily grounded cues, specifically pixel-space actions and explicit spatial-temporal alignment. Even when bypassing this zero-shot barrier via few-shot adaptation, the forced appearance recovery triggers catastrophic forgetting of seen embodiments. Together, these failures expose a critical inability to apply learned physical dynamics to novel visual appearances, highlighting that achieving true cross-embodiment generalization requires architectural innovations that decouple visual appearance from underlying physical dynamics.

世界模型机器人泛化能力跨形态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。