让数字孪生能模拟人手动作交互,生成真实连贯的动态视频。
Dexterous World Models
- 基于静态场景和手部动作序列,生成具物理合理性的交互视频。
- 支持抓取、开合、移动等动作,保持相机与场景一致性。
- 适合虚拟仿真、人机交互研究者使用。
近期3D重建进展使从日常环境创建逼真数字孪生变得容易,但现有数字孪生仍以静态为主,仅限于导航与视角合成,缺乏具身交互能力。为此,我们提出灵巧世界模型(Dexterous World Model, DWM),一种场景-动作条件化的视频扩散框架,可建模灵巧人类动作如何引起静态3D场景的动态变化。给定静态3D场景渲染图与第一人称手部运动序列,DWM生成时间连贯、符合物理规律的人-场景交互视频。该方法通过两种条件实现:(1)沿指定相机轨迹的静态场景渲染,确保空间一致性;(2)编码几何与运动信息的第一人称手部网格渲染,直接建模动作条件下的动态行为。为训练DWM,我们构建了一个混合交互视频数据集:合成的第一人称交互提供完整对齐监督,用于联合运动与操作学习;固定摄像头的真实视频则引入多样且真实的物体动态。实验表明,DWM能生成逼真且物理合理的交互行为,如抓取、打开、移动物体,同时保持相机与场景一致。该框架为基于视频扩散的交互式数字孪生迈出第一步,支持从第一人称动作进行具身仿真。
原文摘要 · Abstract (English)
Recent progress in 3D reconstruction has made it easy to create realistic digital twins from everyday environments. However, current digital twins remain largely static and are limited to navigation and view synthesis without embodied interactivity. To bridge this gap, we introduce Dexterous World Model (DWM), a scene-action-conditioned video diffusion framework that models how dexterous human actions induce dynamic changes in static 3D scenes. Given a static 3D scene rendering and an egocentric hand motion sequence, DWM generates temporally coherent videos depicting plausible human-scene interactions. Our approach conditions video generation on (1) static scene renderings following a specified camera trajectory to ensure spatial consistency, and (2) egocentric hand mesh renderings that encode both geometry and motion cues to model action-conditioned dynamics directly. To train DWM, we construct a hybrid interaction video dataset. Synthetic egocentric interactions provide fully aligned supervision for joint locomotion and manipulation learning, while fixed-camera real-world videos contribute diverse and realistic object dynamics. Experiments demonstrate that DWM enables realistic and physically plausible interactions, such as grasping, opening, and moving objects, while maintaining camera and scene consistency. This framework represents a first step toward video diffusion-based interactive digital twins and enables embodied simulation from egocentric actions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。