统一建模文本与人-物交互运动,支持多条件生成
Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction

- 用大语言模型和两个VQ-VAE将异构运动数据转为文本兼容的序列
- 在大规模数据上多任务训练,再微调特定任务,性能显著提升
- 适合做虚拟现实、混合现实中多模态交互生成的研究者
建模4D人-物交互(HOI)是计算机视觉中的关键挑战,也是虚拟与混合现实应用的核心技术。现有方法在特定任务如文本驱动的HOI生成、由物体运动生成人体动作等方面取得进展,但普遍依赖任务专用架构,缺乏统一框架处理多种条件输入。为此,我们提出Uni-HOI:一个联合学习文本、人体运动与物体运动分布的统一框架。通过大语言模型(LLMs)和两个针对运动的向量量化变分自编码器(VQ-VAEs),将异构运动数据转换为适配LLM输入的标记序列,实现三模态的无缝集成与联合建模。采用两阶段训练策略:第一阶段在大规模HOI数据集上进行多任务学习,捕捉三者间的潜在关联;第二阶段针对具体任务微调以进一步提升性能。大量实验表明,Uni-HOI在多个相关任务中表现卓越,包括文本驱动的HOI生成、物体运动驱动的人体动作生成(可选附加文本)、以及人体动作驱动的物体运动预测,均在一个统一框架内完成。
原文摘要 · Abstract (English)
Modeling 4D human-object interaction (HOI) is a compelling challenge in computer vision and an essential technology powering virtual and mixed-reality applications. While existing works have achieved promising results on specific HOI tasks-such as text-conditioned HOI generation and human motion generation from object motion, they typically rely on task-specific architectures and lack a unified framework capable of handling diverse conditional inputs. Building on this, we propose Uni-HOI, a unified framework that learns the joint distribution among text, human motion, and object motion. By leveraging large language models (LLMs) and two motion-specific vector quantized variational autoencoders (VQ-VAEs), we convert heterogeneous motion data into token sequences compatible with LLM inputs, enabling seamless integration and joint modeling of all three modalities. We introduce a two-stage training strategy: the first stage performs multi-task learning on a large-scale HOI dataset to capture the underlying correlations among the three modalities, while the second stage fine-tunes the model on specific tasks to further enhance performance. Extensive experiments demonstrate that Uni-HOI achieves remarkable performances on multiple HOI-related tasks including text-driven HOI generation, object motion-driven human motion generation (optionally with text) and human motion-driven object motion prediction within a unified framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。