arXiv:2606.05160cs.RO2026-06被引 4

用虚拟3D场景和视频模型生成机器人操作数据,免去真实拍摄和遥控。

GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors

论文配图:GRAIL: Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors
图 1 · 摘自论文原文
  • 用已知3D结构的虚拟环境生成带物理约束的机器人动作数据。
  • 生成超2万条交互序列,实测抓取成功率84%,爬楼梯90%。
  • 适合做仿人机器人操控、无需真实采集数据的研究者。

实现仿人机器人多任务操作需大量兼容机器人的示范数据,涵盖不同物体、全身运动与场景结构,但远程操控与动捕难以规模化,因每次采集依赖实体布置、传感器演员与机器人运行。本文提出GRAIL,一种全程虚拟的数据生成流程:通过组合3D资产、模拟器可用场景及视频基础模型(VFMs)先验,合成交互数据,无需重建物理环境或遥控机器人。GRAIL不从无约束野外视频恢复,而是基于预先确定的3D配置——包括物体几何、相机参数、度量尺度、环境深度与适配机器人尺寸的角色——进行视频生成并复用于重建。这一优势设定更利于4维恢复,支持基于模型的对象跟踪、人体运动估计与交互感知优化,显著降低深度模糊与形态不匹配。我们将恢复的动作重定向至仿人机器人,并训练互补的任务通用追踪器:一个对象感知的隐空间适配器用于操作,一个场景感知追踪器用于地形穿越。GRAIL生成超过20,000个序列,涵盖抓取、操作、坐下与地形穿越。仅使用生成数据,通过端到端仿真到现实策略训练,部署于Unitree G1仿人机器人,在多样物体抓取任务中实现84%真实世界成功率,楼梯攀爬达90%。

原文摘要 · Abstract (English)

Scaling humanoid loco-manipulation requires robot-compatible demonstrations across diverse objects, whole-body motions, and scene geometries, but teleoperation and motion capture are difficult to scale because each collection depends on physical setups, instrumented actors, and robot operation. We present GRAIL, a digital generation pipeline that remains fully virtual until deployment: it composes 3D assets, simulator-ready scenes, and priors from video foundation models (VFMs) to synthesize interactions without rebuilding physical environments or teleoperating the robot. Rather than reconstructing unconstrained in-the-wild videos, GRAIL starts from fully specified 3D configurations in which object geometry, camera parameters, metric scale, environment depth, and a robot-proportioned character are known before video generation and reused during reconstruction. This privileged setup better conditions 4D recovery, allowing model-based object tracking, human motion estimation, and interaction-aware optimization to reconstruct metric 4D human-object interaction (HOI) trajectories with reduced depth ambiguity and morphology mismatch. We retarget the recovered motions to a humanoid robot and train complementary task-general trackers: an object-aware latent adaptor for manipulation and a scene-aware tracker for terrain traversal. GRAIL produces over 20,000 sequences spanning pick-up, object manipulation, sitting, and terrain traversal. Using only GRAIL-generated data, we train egocentric visual policies through a sim-to-real pipeline and deploy them on a Unitree G1 humanoid, achieving 84\% real-world success on diverse object pick-up and 90\% success on stair-climbing.

仿人机器人动作生成虚拟数据视觉控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。