构建百万小时多模态游戏数据集,助力真实世界智能体训练
PLAICraft: Large-Scale Time-Aligned Vision-Speech-Action Dataset for Embodied AI
- 采集10,000+玩家在Minecraft中五维同步数据:视频、音频、键鼠操作
- 超10,000小时高质量实时交互数据,支持毫秒级时间对齐分析
- 适用于研究语言理解、空间感知与长期记忆的具身智能模型评估
深度生成建模的发展使训练人类水平的具身智能体成为可能,但受限于缺乏大规模、实时、多模态且具备社会互动性的数据集。为此,我们提出PLAICraft,一个新型数据采集平台与数据集,记录了来自全球超过10,000名玩家在多人Minecraft中的交互行为,涵盖视频、游戏输出音频、麦克风输入音频、鼠标和键盘动作五个时间对齐模态,每项均以毫秒级精度记录。该数据集包含超过10,000小时的游戏时长,支持对同步具身行为的深入研究。同时,我们提供评估套件,用于评测模型在物体识别、空间意识、语言定位与长期记忆方面的能力。PLAICraft为训练和评估实时、流畅、有目的性的智能体提供了新路径,推动真正具身人工智能的发展。
原文摘要 · Abstract (English)
Advances in deep generative modeling have made it increasingly plausible to train human-level embodied agents. Yet progress has been limited by the absence of large-scale, real-time, multi-modal, and socially interactive datasets that reflect the sensory-motor complexity of natural environments. To address this, we present PLAICraft, a novel data collection platform and dataset capturing multiplayer Minecraft interactions across five time-aligned modalities: video, game output audio, microphone input audio, mouse, and keyboard actions. Each modality is logged with millisecond time precision, enabling the study of synchronous, embodied behaviour in a rich, open-ended world. The dataset comprises over 10,000 hours of gameplay from more than 10,000 global participants. Alongside the dataset, we provide an evaluation suite for benchmarking model capabilities in object recognition, spatial awareness, language grounding, and long-term memory. PLAICraft opens a path toward training and evaluating agents that act fluently and purposefully in real time, paving the way for truly embodied artificial intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。