让虚拟世界实时响应用户手部和头部动作,实现更自然的沉浸式交互。
Generated Reality: Human-centric World Simulation using Interactive Video Generation with Hand and Camera Control
- 用头部和手部关节姿态作为控制信号,实现精细动作引导视频生成。
- 用户任务完成率提升,感知控制感显著高于基线方法。
- 适合元宇宙、VR交互等需要身体动作反馈的应用场景。
扩展现实(XR)需要能够响应用户真实世界运动的生成模型,但现有视频世界模型仅支持文本或键盘等粗粒度输入,难以实现具身交互。本文提出一种以人为中心的视频世界模型,同时以追踪的头部姿态和手部关节姿态为条件。评估了现有的扩散变换器条件策略,并提出一种有效的3D头部与手部控制机制,支持灵巧的手物交互。训练一个双向视频扩散教师模型,并将其蒸馏为因果性、可交互的系统,生成第一人称视角的虚拟环境。通过人类受试者实验验证,该生成现实系统在任务表现上有所提升,且用户对操作行为的控制感显著高于相关基线。
原文摘要 · Abstract (English)
Extended reality (XR) demands generative models that respond to users' tracked real-world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied interaction. We introduce a human-centric video world model that is conditioned on both tracked head pose and joint-level hand poses. For this purpose, we evaluate existing diffusion transformer conditioning strategies and propose an effective mechanism for 3D head and hand control, enabling dexterous hand--object interactions. We train a bidirectional video diffusion model teacher using this strategy and distill it into a causal, interactive system that generates egocentric virtual environments. We evaluate this generated reality system with human subjects and demonstrate improved task performance as well as a significantly higher level of perceived amount of control over the performed actions compared with relevant baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。