仅用单个头戴摄像头生成逼真人体动作,支持实时推理。
HMD^2: Environment-aware Motion Generation from Single Egocentric Head-Mounted Device
- 结合视觉SLAM与图像嵌入,融合解析与学习特征重建头部运动
- 采用多模态扩散模型生成连贯动作,0.17秒延迟实现在线推理
- 在200小时复杂环境数据上表现稳健,适合真实场景应用
本文研究如何仅使用配备外向彩色摄像头和视觉SLAM能力的单个头戴设备,生成逼真全身人体动作。为解决该设置下的模糊性问题,我们提出HMD^2,一种在动作重建与生成间取得平衡的新系统。从重建角度,它最大化利用摄像头流,提取包括头部运动、SLAM点云和图像嵌入在内的分析与学习特征。在生成方面,HMD^2采用基于Transformer骨干的多模态条件动作扩散模型,保持动作的时间一致性,并通过自回归填充实现低延迟在线推理(0.17秒)。实验表明,该系统在超过200小时、涵盖复杂室内外环境的多样化数据集上表现有效且鲁棒。
原文摘要 · Abstract (English)
This paper investigates the generation of realistic full-body human motion using a single head-mounted device with an outward-facing color camera and the ability to perform visual SLAM. To address the ambiguity of this setup, we present HMD^2, a novel system that balances motion reconstruction and generation. From a reconstruction standpoint, it aims to maximally utilize the camera streams to produce both analytical and learned features, including head motion, SLAM point cloud, and image embeddings. On the generative front, HMD^2 employs a multi-modal conditional motion diffusion model with a Transformer backbone to maintain temporal coherence of generated motions, and utilizes autoregressive inpainting to facilitate online motion inference with minimal latency (0.17 seconds). We show that our system provides an effective and robust solution that scales to a diverse dataset of over 200 hours of motion in complex indoor and outdoor environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。