arXiv:2607.28625cs.CV2026-07

构建可同步采集多模态人因数据的环境捕捉系统,支持具身智能研究。

ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine

论文配图:ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
图 1 · 摘自论文原文
  • 设计双尺度环境捕捉引擎,同步记录第一视角与多视角视频、全身动作等多模态数据。
  • 建成150小时、1700万帧的ACE-Data-0数据集,涵盖200类任务与7.5万次交互事件。
  • 适合具身智能、模仿学习、视觉-语言-行动系统等方向的研究者使用。

具身智能面临核心数据瓶颈:模型需同时捕捉第一人称感知、全身运动、灵巧操作、物体状态、声音和触觉随时间演化的完整感知-行动闭环。现有数据集将体验割裂于不同视角、模态或空间尺度,导致信息不全。本文提出环境捕捉引擎(ACE),将真实家居环境转化为时空同步、空间校准的录制平台。其表尺度配置解析手物操作,房间尺度配置捕捉全身运动与跨空间交互。ACE统一记录第一人称与多视角外视角视频、全身与关节手部动作、物体几何与6-DoF轨迹、音频及触觉信号。基于此构建了ACE-Data-0数据集,包含150小时、1700万帧视频,覆盖200类任务,由50名参与者在2个环境中完成,共75,000次交互。数据涵盖原子级操作、长时程家庭活动链与人-场景交互,通过目标级而非步骤级指令保留自然行为变异。我们进一步提出分层评估基准,从信号到场景组件再到交互。对前沿方法的评估揭示其在接触、遮挡、自我运动及长时序条件下存在显著差距。ACE-Data-0提供同步的人类示范数据,附带对齐的感知、运动与接触监督,为模仿学习、世界模型、视觉-语言-行动系统及具身智能提供可扩展基础。

原文摘要 · Abstract (English)

Embodied intelligence faces a fundamental data bottleneck. Models must capture how first-person perception, whole-body motion, dexterous manipulation, object state, sound, and touch evolve together as humans pursue goals over time. Existing datasets fragment this experience across viewpoints, modalities, or spatial scales, leaving the full perception-action loop only partially observed. We introduce the Ambient Capture Engine (ACE), a human-centric data engine that transforms real home environments into spatially calibrated, temporally synchronized recording studios. ACE operates at two complementary scales: a table-scale configuration resolves hand-object manipulation, while a room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. ACE records egocentric and multi-view exocentric video, full-body and articulated hand motion, object geometry and 6-DoF trajectories, audio, and tactile signals as a unified multisensory stream. Using ACE, we build ACE-Data-0, comprising 150 hours and 17M video frames across 200 task categories, performed by 50 participants in 2 environments, for a total of 75,000 interaction episodes. The dataset spans atomic manipulation, long-horizon chains of household activities, and human-scene interaction, while preserving natural behavioral variation through goal-level rather than step-by-step instructions. We further introduce a hierarchical benchmark that progresses from signals to scene components and then to interactions. Evaluations of state-of-the-art methods expose substantial gaps under contact, occlusion, egomotion, and long temporal horizons. ACE-Data-0 provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision, offering a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI.

具身智能多模态数据模仿学习数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。