用部件功能引导生成更真实的零样本4D人物交互动作
HOI-PAGE: Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance
- 通过大模型构建部件功能图,指导交互生成
- 支持多对象多人复杂交互,生成动作更真实
- 适合需要高保真交互生成的虚拟场景应用
我们提出HOI-PAGE,一种基于部件功能引导的零样本4D人-物交互生成方法。与以往关注整体身体-物体运动的工作不同,该方法利用大语言模型显式推理交互的部件级机制,构建结构化的部件功能图(PAG)作为高层交互骨架。其三阶段合成流程包括:将3D物体分解为语义部件;根据文本提示生成参考交互视频以提取基于部件的运动约束;最终优化生成满足部件接触约束且模仿参考动态的4D交互运动序列。大量实验表明,该方法在零样本4D HOI生成中显著提升了真实感和文本对齐度,可灵活生成复杂的多物体或多角色交互序列。
原文摘要 · Abstract (English)
We present HOI-PAGE, a new approach that prioritizes part-level affordance reasoning to generate high-fidelity 4D human-object interactions (HOIs) from text prompts in a zero-shot fashion. In contrast to prior works that focus on global, whole body-object motion synthesis, our approach explicitly reasons about the underlying part-level mechanics of interactions using large language models (LLMs). We capture this reasoning in a structured part affordance graph (PAG) representation, serving as a high-level interaction scaffolding to guide a three-stage synthesis: first, decomposing input 3D objects into semantic parts; then, generating reference HOI videos from text prompts to extract part-based motion constraints; and finally, optimizing for 4D HOI motion sequences that mimic the reference dynamics while satisfying part-level contact constraints. Extensive experiments show that our approach is flexible and capable of generating complex multi-object or multi-person interaction sequences, with significantly improved realism and text alignment for zero-shot 4D HOI generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。