开源人体视角操作数据集,支持机器人学习全流程
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

- 用手机连续采集2000小时自然环境操作视频
- 提供动作分割、手部姿态、相机轨迹等结构化标注
- 适配视觉语言模型与世界模型训练,支持人机迁移
人体视角操作视频为具身智能提供了可扩展的监督信号,但现有资源极少同时具备低成本连续采集、操作级结构化标注和可复用的机器人学习工具链。我们提出 Open-AoE,一个面向社区的开源人体视角操作数据集与全流程工具链,覆盖从手机采集到模型训练的全环节。首个版本包含约2000小时操作视频,由500多名贡献者使用400多部智能手机在自然环境中采集。数据集提供文本标注、基于MANO的手部姿态、相机轨迹及时间定位的原子动作。Open-AoE还包含数据处理流水线,通过动作时序分割、语义标注、手部重建和相机轨迹恢复,将原始录制转化为结构化样本。此外,下游工具链支持可视化、跨具身性重定向、模型专用数据转换以及针对VLA策略、WAMs和世界模型的训练方案。通过整合可扩展采集、结构化处理与下游适配,Open-AoE降低了数据贡献与复用门槛,为具身模型训练、人机迁移与世界建模提供实用开放基础设施。
原文摘要 · Abstract (English)
Egocentric videos of human manipulation provide scalable supervision for embodied intelligence, yet existing resources rarely combine low-cost continuous capture, manipulation-level structured annotations, and reusable tools for robot learning. We present Open-AoE, an open, community-oriented egocentric manipulation dataset and toolchain spanning the full pipeline from smartphone capture to model training. Its first release contains approximately 2,000 hours of manipulation video collected in natural environments by 500+ contributors using 400+ smartphones. The dataset provides text annotations, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. Open-AoE further includes a data processing pipeline that transforms raw recordings into structured samples through temporal action segmentation, semantic annotation, hand reconstruction, and camera trajectory reconstruction. Meanwhile, we provide a separate downstream toolchain supports visualization, cross-embodiment retargeting, model-specific data conversion, and training recipes for VLA policies, WAMs, and World Models. By integrating scalable capture, structured processing, and downstream adaptation, Open-AoE reduces the barriers to both data contribution and reuse, providing practical open infrastructure for embodied model training, human-to-robot transfer, and world modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。