arXiv:2608.15141cs.CV2026-08

HOIMask用离散化掩码建模生成人物交互动作,更真实稳定。

HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation

论文配图:HOIMask: Towards Generative Masked Modeling for Human Object Interaction Generation
图 1 · 摘自论文原文
  • 将动作与接触信号转为离散2D令牌图,保留时空结构
  • 通过掩码建模提升交互一致性,生成动作更连贯
  • 适合需要高精度交互动作生成的研究者

基于扩散的方法在人-物交互(HOI)生成中占主导地位,因其能通过关键接触信号引导扩散过程。然而,迭代去噪中的误差累积常导致高伪影和交互质量不稳定。本文提出HOIMask,首个面向离散空间的生成式掩码建模框架,用于建模人-物交互运动。首先,通过人-物向量量化(HOI VQ)将运动序列与接触感知信号编码为离散2D人体与物体令牌图,保留了超越传统1D表示的精细时空结构。在此基础上,采用生成式掩码建模框架,利用专为复杂时空与交互依赖设计的Transformer架构,联合捕捉人-物交互动态。为生成更连贯且物理合理的动作,推理阶段引入新颖的接触感知重建引导,融合接触信号优化交互令牌,强化生成动作的时空一致性。凭借精心设计的动作交互令牌、专用架构与引导策略,HOIMask超越现有扩散方法,生成更真实、语义对齐的HOI动作。

原文摘要 · Abstract (English)

Diffusion-based methods have dominated the HOI generation, as they enable critical contact fusions or signals to guide the diffusion process. However, they often result in high artifacts and unstable interaction quality due to error accumulation during iterative denoising. In this work, we propose HOIMask, the first generative masked framework for modeling HOI motion in discrete space. HOIMask first encodes both motion sequences and contact-aware signals into discrete 2D human and object token maps via HOI Vector Quantization (VQ), preserving fine-grained spatial-temporal structure beyond conventional 1D representations. On this basis, a generative masked modeling framework is employed to jointly capture human-object interaction dynamics, leveraging a transformer architecture designed to model complex spatial-temporal and interaction dependencies. To generate more coherent and physically plausible motions, we further introduce a novel contact-aware reconstruction guidance in discrete space during inference, which fuses contact signals to optimize HOI tokens that forces the generated motion with higher spatio-temporal consistency. With craftily designed motion interaction tokens, dedicated architecture and guidance strategy, HOIMask outperforms state-of-the-art diffusion-based methods, generating more realistic and semantically aligned HOI motions. Please refer to https://jyhflash.github.io/HOIMask/ for more results.

人-物交互生成模型扩散模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。