arXiv:2511.20446cs.CV2025-11NeurIPS被引 2

让机器根据文字生成多人与物体的自然互动场景。

Learning to Generate Human-Human-Object Interactions from Textual Descriptions

  • 用扩散模型联合建模人-物和人-人交互,生成完整互动过程。
  • 在真实性和多样性上优于仅处理单人的现有方法。
  • 适合需要多角色动态交互生成的虚拟仿真与视频创作。

人类之间的互动,包括人际距离、空间布局和运动模式,在不同情境中差异显著。为使机器理解这种复杂且依赖上下文的行为,需建模多人与环境的关系。本文提出新研究问题:建模两人共同参与物体交互时的关联行为,称为人-人-物交互(HHOIs)。针对缺乏专用数据集的问题,我们构建了一个新采集的HHOIs数据集,并利用图像生成模型合成相关数据。作为中间步骤,从HHOIs中提取个体的人-物交互(HOIs)和人-人交互(HHIs),并基于得分驱动的扩散模型训练文本到HOI和文本到HHI的生成模型。最终提出统一生成框架,通过一次采样即可合成完整的HHOIs。该方法扩展至多人群体场景,支持超过两人参与的交互生成。实验表明,该方法能生成符合文本描述的真实感强的HHOIs,性能超越仅关注单人的先前方法。此外,我们还将框架应用于含物体的多人群体动作生成任务。

原文摘要 · Abstract (English)

The way humans interact with each other, including interpersonal distances, spatial configuration, and motion, varies significantly across different situations. To enable machines to understand such complex, context-dependent behaviors, it is essential to model multiple people in relation to the surrounding scene context. In this paper, we present a novel research problem to model the correlations between two people engaged in a shared interaction involving an object. We refer to this formulation as Human-Human-Object Interactions (HHOIs). To overcome the lack of dedicated datasets for HHOIs, we present a newly captured HHOIs dataset and a method to synthesize HHOI data by leveraging image generative models. As an intermediary, we obtain individual human-object interaction (HOIs) and human-human interaction (HHIs) from the HHOIs, and with these data, we train an text-to-HOI and text-to-HHI model using score-based diffusion model. Finally, we present a unified generative framework that integrates the two individual model, capable of synthesizing complete HHOIs in a single advanced sampling process. Our method extends HHOI generation to multi-human settings, enabling interactions involving more than two individuals. Experimental results show that our method generates realistic HHOIs conditioned on textual descriptions, outperforming previous approaches that focus only on single-human HOIs. Furthermore, we introduce multi-human motion generation involving objects as an application of our framework.

交互生成扩散模型多角色文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。