统一生成与编辑人体物体重叠关系,支持多种控制条件。
OneHOI: Unifying Human-Object Interaction Generation and Editing

- 用共享结构化表示统一建模生成与编辑过程。
- 在HOI-Edit-44K数据集上达到当前最佳性能。
- 适合需要灵活控制交互场景的研究者和开发者。
人体-物体交互(HOI)建模捕捉人类如何作用于物体,通常以<人, 动作, 物体>三元组形式表达。现有方法分为两类:一类是基于结构化三元组和布局的场景生成,但无法整合包含混合条件(如交互与仅物体实体)的情况;另一类是通过文本修改交互,却难以解耦姿态与物理接触,且难以扩展到多交互场景。我们提出OneHOI,一个统一的扩散Transformer框架,将HOI生成与编辑合并为单一条件去噪过程,由共享的结构化交互表示驱动。核心是关系扩散Transformer(R-DiT),通过角色与实例感知的HOI token、基于布局的空间动作定位、结构化HOI注意力以强制交互拓扑,并采用HOI RoPE分离多交互场景。在我们的HOI-Edit-44K数据集上联合训练,并融合其他HOI与物体中心数据集,OneHOI支持布局引导、无布局、任意掩码及混合条件控制,在生成与编辑任务上均达到最先进水平。代码已公开。
原文摘要 · Abstract (English)
Human-Object Interaction (HOI) modelling captures how humans act upon and relate to objects, typically expressed as <person, action, object> triplets. Existing approaches split into two disjoint families: HOI generation synthesises scenes from structured triplets and layout, but fails to integrate mixed conditions like HOI and object-only entities; and HOI editing modifies interactions via text, yet struggles to decouple pose from physical contact and scale to multiple interactions. We introduce OneHOI, a unified diffusion transformer framework that consolidates HOI generation and editing into a single conditional denoising process driven by shared structured interaction representations. At its core, the Relational Diffusion Transformer (R-DiT) models verb-mediated relations through role- and instance-aware HOI tokens, layout-based spatial Action Grounding, a Structured HOI Attention to enforce interaction topology, and HOI RoPE to disentangle multi-HOI scenes. Trained jointly with modality dropout on our HOI-Edit-44K, along with HOI and object-centric datasets, OneHOI supports layout-guided, layout-free, arbitrary-mask, and mixed-condition control, achieving state-of-the-art results across both HOI generation and editing. Code is available at https://jiuntian.github.io/OneHOI/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。