arXiv:2606.04463cs.RO2026-06

OSCAR让机器人在虚拟世界中精准执行动作并跨平台评估策略。

OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics

论文配图:OSCAR: Omni-Embodiment Action-Conditioned World Model for Robotics
图 1 · 摘自论文原文
  • 用统一骨骼渲染作为动作条件,适配不同机器人和人手。
  • 在单张GH200 GPU上微调后,动作跟随与画面质量显著提升。
  • 可替代真实测试,为机器人策略提供高相关性虚拟评估。

我们提出OSCAR,一种高精度动作条件视频世界模型,能跨不同机器人形态泛化并用于机器人策略评估。现有视频世界模型在真实机器人评估中面临三大挑战:训练数据场景多样性不足、动作跟随不精确、跨形态泛化能力差。为此,我们构建大规模标准化数据流水线,整合并清洗机器人及第一人称人类数据,生成覆盖多样任务、场景、动作和机器人形态的联合训练数据集。采用2D运动学骨架渲染作为统一条件表示,实现对不同机械臂甚至人手的通用建模。在单张GH200 GPU上微调Cosmos-Predict2.5-2B模型,相比基线方法(更大模型或更多GPU),在动作跟随、外观质量和运动一致性上均有显著提升。进一步将OSCAR部署于RoboArena评估机器人策略,实验表明其虚拟评估结果与真实世界表现高度相关,为未来纯虚拟环境下机器人策略评估铺平道路。

原文摘要 · Abstract (English)

We present OSCAR, a precise action-conditioned video world model that generalizes across different robot embodiments and enables robot policy evaluation. Existing video world models face three main challenges for real-world robot evaluation: limited scenario diversity in current robot training datasets, imprecise action following, and poor generalization across embodiments for broad adoption. We tackle these challenges from two perspectives. At its core is a large-scale standardized data pipeline that curates, filters, and deduplicates broad robotics and egocentric human datasets, yielding a clean joint-training dataset that spans diverse tasks, scenarios, actions, and robot embodiments. To condition the video model, we adopt 2D kinematic skeleton rendering as a unified conditioning representation that generalizes across different robot arms or even human hands. We finetune the Cosmos-Predict2.5-2B model on a single GH200 GPU. Our model achieves significant improvement on action following, appearance quality, and motion consistency, compared to existing baselines, which either have a much larger model size or require more GPUs. We further deploy OSCAR to evaluate robot policies from RoboArena. Extensive experiments demonstrate the significant correlation between our virtual policy evaluation in OSCAR and real-world evaluation, paving the way for the future where robot policies can be purely evaluated in virtual generated worlds.

机器人世界模型动作条件虚拟评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。