用Transformer统一生成全身抓取动作,更真实稳定。
A Unified Transformer-Based Framework with Pretraining For Whole Body Grasping Motion Generation
- 分三阶段:先生成抓取姿态,再补全时间连续动作,最后提升关节分辨率。
- 在GRAB数据集上比现有方法更连贯、稳定且视觉真实。
- 预训练提升数据效率,适合想做人体动作生成的开发者。
我们提出一种基于Transformer的全身抓取动作生成框架,同时解决抓取姿态生成与动作补全问题,实现逼真稳定的物体交互。该流程包含三个阶段:抓取姿态生成(全身体态生成)、时间补全(保持动作连续性)以及一个用于将下采样关节还原为高分辨率标记的LiftUp Transformer。为应对手物交互数据稀缺问题,我们在大规模多样运动数据集上引入高效通用预训练阶段,获得可迁移至抓取任务的鲁棒时空表征。在GRAB数据集上的实验表明,该方法在连贯性、稳定性与视觉真实性方面优于现有最先进基线。模块化设计也便于拓展至其他人体动作应用。
原文摘要 · Abstract (English)
Accepted in the ICIP 2025 We present a novel transformer-based framework for whole-body grasping that addresses both pose generation and motion infilling, enabling realistic and stable object interactions. Our pipeline comprises three stages: Grasp Pose Generation for full-body grasp generation, Temporal Infilling for smooth motion continuity, and a LiftUp Transformer that refines downsampled joints back to high-resolution markers. To overcome the scarcity of hand-object interaction data, we introduce a data-efficient Generalized Pretraining stage on large, diverse motion datasets, yielding robust spatio-temporal representations transferable to grasping tasks. Experiments on the GRAB dataset show that our method outperforms state-of-the-art baselines in terms of coherence, stability, and visual realism. The modular design also supports easy adaptation to other human-motion applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。