arXiv:2509.05513cs.CVcs.AI2025-09被引 10

构建大规模多模态第一视角操作数据集,支持精细手部动作学习。

OpenEgo: A Large-Scale Multimodal Egocentric Dataset for Dexterous Manipulation

  • 统一手部姿态标注格式,提供时间对齐的动作原子描述。
  • 覆盖290项操作任务,总计1107小时视频数据,来自6个公开数据集。
  • 适合研究视觉-语言-动作协同学习与灵巧操作的复现工作。

第一视角人类视频为模仿学习提供了可扩展的示范数据,但现有数据集通常缺乏细粒度且时间精准的动作描述或灵巧手部标注。我们提出OpenEgo,一个包含标准化手部姿态标注和意图对齐动作原子的多模态第一视角操作数据集。OpenEgo总计1107小时,涵盖6个公开数据集,覆盖290项操作任务、600多个环境。我们统一了手部姿态布局,并提供描述性、带时间戳的动作原子。为验证其有效性,我们训练了语言条件下的模仿学习策略以预测灵巧手部轨迹。OpenEgo旨在降低从第一视角视频学习灵巧操作的门槛,并支持视觉-语言-动作学习领域的可复现研究。所有资源与使用说明将发布于www.openegocentric.com。

原文摘要 · Abstract (English)

Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal egocentric manipulation dataset with standardized hand-pose annotations and intention-aligned action primitives. OpenEgo totals 1107 hours across six public datasets, covering 290 manipulation tasks in 600+ environments. We unify hand-pose layouts and provide descriptive, timestamped action primitives. To validate its utility, we train language-conditioned imitation-learning policies to predict dexterous hand trajectories. OpenEgo is designed to lower the barrier to learning dexterous manipulation from egocentric video and to support reproducible research in vision-language-action learning. All resources and instructions will be released at www.openegocentric.com.

第一视角灵巧操作多模态数据集模仿学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。