用自然语言生成新物体的复杂操作动作,突破传统方法局限。
OpenHOI: Open-World Hand-Object Interaction Synthesis with Multimodal Large Language Model
- 结合大模型理解指令并分解任务,定位交互区域。
- 通过扩散模型生成合理动作,物理优化减少穿透。
- 支持未知物体和复杂语言指令,适合机器人与VR应用。
理解与合成真实的3D手物交互(HOI)对沉浸式AR/VR和灵巧机器人至关重要。现有方法在封闭集物体和预定义任务上表现良好,但难以处理未见物体或开放词汇指令。我们提出OpenHOI,首个面向开放世界的手物交互合成框架,可基于自由文本命令生成针对新物体的长时序操作序列。方法融合微调后的3D多模态大语言模型(MLLM),实现交互区域(如把手、按钮)的精准定位与复杂指令(如“找水瓶并喝一口”)的语义任务分解。为生成物理合理的交互,提出基于可及性驱动的扩散模型,并引入无需训练的物理精修阶段,有效减少穿透并优化可及性对齐。在多种场景下的评估表明,OpenHOI在泛化到新物体类别、多阶段任务及复杂语言指令方面优于现有最先进方法。
原文摘要 · Abstract (English)
Understanding and synthesizing realistic 3D hand-object interactions (HOI) is critical for applications ranging from immersive AR/VR to dexterous robotics. Existing methods struggle with generalization, performing well on closed-set objects and predefined tasks but failing to handle unseen objects or open-vocabulary instructions. We introduce OpenHOI, the first framework for open-world HOI synthesis, capable of generating long-horizon manipulation sequences for novel objects guided by free-form language commands. Our approach integrates a 3D Multimodal Large Language Model (MLLM) fine-tuned for joint affordance grounding and semantic task decomposition, enabling precise localization of interaction regions (e.g., handles, buttons) and breakdown of complex instructions (e.g., "Find a water bottle and take a sip") into executable sub-tasks. To synthesize physically plausible interactions, we propose an affordance-driven diffusion model paired with a training-free physics refinement stage that minimizes penetration and optimizes affordance alignment. Evaluations across diverse scenarios demonstrate OpenHOI's superiority over state-of-the-art methods in generalizing to novel object categories, multi-stage tasks, and complex language instructions. Our project page at \href{https://openhoi.github.io}
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。