PAM统一生成手姿、外观和动态,实现从仿真到现实的可控手物交互视频生成。
PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
- 构建统一框架,联合建模手部姿态、外观与运动,支持端到端生成。
- 在DexYCB上实现29.13的FVD(优于基线9.7)和19.37mm的MPJPE(优于基线10.68)。
- 生成480x720高清视频,且合成数据可提升真实数据不足时的手姿估计性能。
手物交互(HOI)重建与合成正成为具身智能与AR/VR的核心。然而现有研究仍分裂为三类:仅预测姿态但不生成像素;基于单图生成外观但缺乏动态;需完整姿态序列与首帧真值输入,无法实现仿真到现实部署。受Joo等(2018)启发,我们提出统一的姿势-外观-运动引擎PAM,用于可控的HOI视频生成。在DexYCB上,我们的方法获得29.13的FVD(对比InterDyn的38.83),以及19.37mm的MPJPE(对比CosHand的30.05),同时生成480x720分辨率视频,高于256x256与256x384基线。在OAKINK2上,全条件模型将FVD从68.76降至46.31。消融实验表明,结合深度、分割与关键点输入效果最佳。此外,使用3,400个合成视频(207,000帧)增强训练,仅用50%真实数据的SimpleHand模型即可达到100%真实数据基准性能。
原文摘要 · Abstract (English)
Hand-object interaction (HOI) reconstruction and synthesis are becoming central to embodied AI and AR/VR. Yet, despite rapid progress, existing HOI generation research remains fragmented across three disjoint tracks: (1) pose-only synthesis that predicts MANO trajectories without producing pixels; (2) single-image HOI generation that hallucinates appearance from masks or 2D cues but lacks dynamics; and (3) video generation methods that require both the entire pose sequence and the ground-truth first frame as inputs, preventing true sim-to-real deployment. Inspired by the philosophy of Joo et al. (2018), we think that HOI generation requires a unified engine that brings together pose, appearance, and motion within one coherent framework. Thus we introduce PAM: a Pose-Appearance-Motion Engine for controllable HOI video generation. The performance of our engine is validated by: (1) On DexYCB, we obtain an FVD of 29.13 (vs. 38.83 for InterDyn), and MPJPE of 19.37 mm (vs. 30.05 mm for CosHand), while generating higher-resolution 480x720 videos compared to 256x256 and 256x384 baselines. (2) On OAKINK2, our full multi-condition model improves FVD from 68.76 to 46.31. (3) An ablation over input conditions on DexYCB shows that combining depth, segmentation, and keypoints consistently yields the best results. (4) For a downstream hand pose estimation task using SimpleHand, augmenting training with 3,400 synthetic videos (207k frames) allows a model trained on only 50% of the real data plus our synthetic data to match the 100% real baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。