arXiv:2602.23205cs.CV2026-02被引 2

用两部移动手机实现户外人体与场景的4D重建,低成本获取真实环境下的动作数据。

EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents

  • 双手机同步拍摄并联合标定,统一重建人体与场景的度量空间坐标。
  • 相比单摄像头,双视角显著降低深度模糊,重建精度更高。
  • 适用于机器人控制、物理动画和单目人-场景重建等具身智能任务。

真实世界中的人类行为自然包含丰富的长期上下文信息,可用于训练具身智能体进行感知、理解与行动。然而,现有采集系统通常依赖昂贵的棚拍环境和可穿戴设备,限制了野外场景化人体动作数据的大规模收集。为此,我们提出EmbodMocap,一种基于两部移动iPhone的便携、低成本数据采集流程。核心思想是联合标定双路RGB-D序列,在统一度量的世界坐标系中同时重建人体与场景。该方法可在日常环境中实现度量尺度一致、场景连贯的捕捉,无需静态相机或标记点,无缝融合人体运动与场景几何。与光学捕获真值对比,双视角设置展现出显著缓解深度模糊的能力,重建对齐效果优于单摄像头或单目模型。基于所采集数据,我们推动三项具身智能任务:单目人-场景重建(微调前馈模型输出度量尺度、世界空间对齐的结果)、基于物理的角色动画(证明数据可用于扩展人-物交互技能与场景感知运动追踪)、机器人运动控制(通过模拟到现实的强化学习训练人形机器人复现视频中的人类动作)。实验验证了本流程的有效性及其对具身人工智能研究的贡献。

原文摘要 · Abstract (English)

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting. However, existing capture systems typically rely on costly studio setups and wearable devices, limiting the large-scale collection of scene-conditioned human motion data in the wild. To address this, we propose EmbodMocap, a portable and affordable data collection pipeline using two moving iPhones. Our key idea is to jointly calibrate dual RGB-D sequences to reconstruct both humans and scenes within a unified metric world coordinate frame. The proposed method allows metric-scale and scene-consistent capture in everyday environments without static cameras or markers, bridging human motion and scene geometry seamlessly. Compared with optical capture ground truth, we demonstrate that the dual-view setting exhibits a remarkable ability to mitigate depth ambiguity, achieving superior alignment and reconstruction performance over single iphone or monocular models. Based on the collected data, we empower three embodied AI tasks: monocular human-scene-reconstruction, where we fine-tune on feedforward models that output metric-scale, world-space aligned humans and scenes; physics-based character animation, where we prove our data could be used to scale human-object interaction skills and scene-aware motion tracking; and robot motion control, where we train a humanoid robot via sim-to-real RL to replicate human motions depicted in videos. Experimental results validate the effectiveness of our pipeline and its contributions towards advancing embodied AI research.

4D重建具身智能移动采集多视角重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。