arXiv:2510.05684cs.AIcs.CV2025-10中稿 · ICLR被引 1

用桌面游戏数据预训练机器人智能体,效果媲美大模型。

D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI

  • 将游戏操作数据统一压缩并标准化,支持大规模预训练
  • 1.3万小时数据训练出10亿参数模型,任务成功率超96%
  • 适合想低成本提升机器人感知与动作能力的研究者

大型语言模型依赖互联网规模文本数据,而具身人工智能仍受限于物理轨迹采集的高昂成本。桌面环境(尤其是游戏)提供丰富且可扩展的感官运动交互,同时保持观察-动作耦合结构,适合作为具身学习的数据源。本文提出D2E框架,证明桌面交互可作为机器人具身任务的有效预训练基础。该框架包含三部分:(1) OWA工具包将多样桌面交互统一为标准格式,实现152倍压缩;(2) Generalist-IDM通过基于时间戳的事件预测,在未见游戏中实现强零样本泛化,支持互联网规模伪标签生成;(3) VAPT将桌面预训练表征迁移至真实世界的操作与导航任务。使用1.3K+小时数据(含259小时人工示范与1000+小时伪标签游戏),10亿参数模型在LIBERO任务上达成96.6%成功率,在CANVAS导航任务上达83.3%,性能媲美或超越7倍更大的模型(如π₀, 3.3B 和 OpenVLA, 7B)。结果表明,数字环境中学习的感官运动原语可有效迁移到现实物理任务,确立桌面预训练为具身人工智能的可行范式。所有资源公开可用:https://worv-ai.github.io/d2e。

原文摘要 · Abstract (English)

Large language models leverage internet-scale text data, yet embodied AI remains constrained by the prohibitive costs of physical trajectory collection. Desktop environments -- particularly gaming -- offer a compelling alternative: they provide rich sensorimotor interactions at scale while maintaining the structured observation-action coupling essential for embodied learning. We present D2E (Desktop to Embodied AI), a framework that demonstrates desktop interactions can serve as an effective pretraining substrate for robotics embodied AI tasks. Unlike prior work that remained domain-specific (e.g., VPT for Minecraft) or kept data proprietary (e.g., SIMA), D2E establishes a complete pipeline from scalable desktop data collection to verified transfer in embodied domains. Our framework comprises three components: (1) the OWA Toolkit that unifies diverse desktop interactions into a standardized format with 152x compression, (2) the Generalist-IDM that achieves strong zero-shot generalization across unseen games through timestamp-based event prediction, enabling internet-scale pseudo-labeling, and (3) VAPT that transfers desktop-pretrained representations to physical manipulation and navigation. Using 1.3K+ hours of data (259 hours of human demonstrations and 1K+ hours of pseudo-labeled gameplay), our 1B-parameter model achieves 96.6% success on LIBERO manipulation and 83.3% on CANVAS navigation, matching or surpassing models up to 7x larger, such as π_{0} (3.3B) and OpenVLA (7B). These results demonstrate that sensorimotor primitives learned from digital interactions transfer effectively to real-world physical tasks, establishing desktop pretraining as a practical paradigm for embodied AI. All resources are publicly available at https://worv-ai.github.io/d2e.

具身智能预训练游戏数据机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。