用隐空间流匹配生成逼真手物交互动作,支持文本控制和跨场景泛化。
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
- 通过融合运动学的变分自编码器构建统一隐空间,捕捉复杂交互动态。
- 采用掩码流匹配模型提升时序连贯性,生成动作更自然流畅。
- 相对初始帧预测物体运动,适合大规模合成数据预训练,泛化性强。
生成真实的3D手物交互(HOI)是计算机视觉与机器人领域的基础挑战,需兼顾时序连贯性与高保真物理合理性。现有方法在学习丰富运动表征和进行时序推理方面仍受限。本文提出HO-Flow框架,从文本和标准3D物体生成真实手物运动序列。首先,利用交互感知变分自编码器,结合手部与物体运动学,将运动序列编码到统一隐空间,以捕获丰富的交互动力学。随后,采用掩码流匹配模型,融合自回归时序推理与连续隐变量生成,提升时序一致性。为进一步增强泛化能力,HO-Flow以初始帧为基准预测物体运动,从而可在大规模合成数据上有效预训练。在GRAB、OakInk和DexYCB三个基准上的实验表明,HO-Flow在物理合理性和动作多样性方面均达到当前最优性能。
原文摘要 · Abstract (English)
Generating realistic 3D hand-object interactions (HOI) is a fundamental challenge in computer vision and robotics, requiring both temporal coherence and high-fidelity physical plausibility. Existing methods remain limited in their ability to learn expressive motion representations for generation and perform temporal reasoning. In this paper, we present HO-Flow, a framework for synthesizing realistic hand-object motion sequences from texts and canoncial 3D objects. HO-Flow first employs an interaction-aware variational autoencoder to encode sequences of hand and object motions into a unified latent manifold by incorporating hand and object kinematics, enabling the representation to capture rich interaction dynamics. It then leverages a masked flow matching model that combines auto-regressive temporal reasoning with continuous latent generation, improving temporal coherence. To further enhance generalization, HO-Flow predicts object motions relative to the initial frame, enabling effective pre-training on large-scale synthetic data. Experiments on the GRAB, OakInk, and DexYCB benchmarks demonstrate that HO-Flow achieves state-of-the-art performance in both physical plausibility and motion diversity for interaction motion synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。