无需训练即可生成自然人物交互,支持任意物体。
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
- 用大语言模型理解指令并引导交互生成
- 通过扩散模型解析物体接触点,实现零样本泛化
- 基于人类反馈优化细节,适合开放集物体交互
近期3D人感知生成取得进展,但现有方法在从文本生成新的人物交互(HOI)方面仍受限,尤其对开放集物体。本文提出一种零样本3D HOI生成框架,不依赖特定数据集训练。利用大语言模型(LLM)推理人-物关系,初始化物体属性并指导优化;采用预训练2D图像扩散模型解析未见物体,提取接触点,避免3D资产知识限制;通过多视角SDS采样生成初始人体姿态;最后引入基于LLM人类级反馈的精细化优化,确保人体与物体间真实3D接触,包括抓握时的手部交互。实验表明,该方法在交互精细度和开放集物体处理能力上优于现有方法。
原文摘要 · Abstract (English)
Recent advances in 3D human-aware generation have made significant progress. However, existing methods still struggle with generating novel Human Object Interaction (HOI) from text, particularly for open-set objects. We identify three main challenges of this task: precise human-object relation reasoning, affordance parsing for any object, and detailed human interaction pose synthesis aligning description and object geometry. In this work, we propose a novel zero-shot 3D HOI generation framework without training on specific datasets, leveraging the knowledge from large-scale pre-trained models. Specifically, the human-object relations are inferred from large language models (LLMs) to initialize object properties and guide the optimization process. Then we utilize a pre-trained 2D image diffusion model to parse unseen objects and extract contact points, avoiding the limitations imposed by existing 3D asset knowledge. The initial human pose is generated by sampling multiple hypotheses through multi-view SDS based on the input text and object geometry. Finally, we introduce a detailed optimization to generate fine-grained, precise, and natural interaction, enforcing realistic 3D contact between the 3D object and the involved body parts, including hands in grasping. This is achieved by distilling human-level feedback from LLMs to capture detailed human-object relations from the text instruction. Extensive experiments validate the effectiveness of our approach compared to prior works, particularly in terms of the fine-grained nature of interactions and the ability to handle open-set 3D objects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。