让3D场景中可移动物体的交互按文本指令精准生成
InteractMove: Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects
- 用3D视觉定位识别目标物体,结合手物协同可接触区域预测
- 支持多类可移动物体、不同交互方式,生成物理合理轨迹
- 适合需要真实感人-物交互生成的研究者与开发者
我们提出一个新任务:在包含可移动物体的3D场景中,根据文本指令生成人-物交互。现有数据集交互类别不足,且通常只考虑静态物体(不改变位置),而收集含可移动物体的数据成本高、难度大。为此,我们构建了InteractMove数据集,通过对齐已有交互数据与场景上下文,具备三大特征:1)包含多个可移动物体的场景,支持带文本控制的交互指令(包括同类别干扰物,需空间与3D场景理解);2)多样化的物体类型与尺寸,支持单手、双手等不同交互模式;3)物理合理的物体操作轨迹。为应对挑战,我们提出新流程:首先使用3D视觉定位模型识别交互目标;接着设计手-物联合亲和力学习,预测不同手部关节与物体部件的接触区域,实现对多样化物体的精准抓取与操作;最后通过局部场景建模与碰撞规避约束优化交互,确保运动物理合理性并避免物体与场景碰撞。大量实验表明,本方法在生成符合文本指令且物理合理的交互方面优于现有方法。
原文摘要 · Abstract (English)
We propose a novel task of text-controlled human object interaction generation in 3D scenes with movable objects. Existing human-scene interaction datasets suffer from insufficient interaction categories and typically only consider interactions with static objects (do not change object positions), and the collection of such datasets with movable objects is difficult and costly. To address this problem, we construct the InteractMove dataset for Movable Human-Object Interaction in 3D Scenes by aligning existing human object interaction data with scene contexts, featuring three key characteristics: 1) scenes containing multiple movable objects with text-controlled interaction specifications (including same-category distractors requiring spatial and 3D scene context understanding), 2) diverse object types and sizes with varied interaction patterns (one-hand, two-hand, etc.), and 3) physically plausible object manipulation trajectories. With the introduction of various movable objects, this task becomes more challenging, as the model needs to identify objects to be interacted with accurately, learn to interact with objects of different sizes and categories, and avoid collisions between movable objects and the scene. To tackle such challenges, we propose a novel pipeline solution. We first use 3D visual grounding models to identify the interaction object. Then, we propose a hand-object joint affordance learning to predict contact regions for different hand joints and object parts, enabling accurate grasping and manipulation of diverse objects. Finally, we optimize interactions with local-scene modeling and collision avoidance constraints, ensuring physically plausible motions and avoiding collisions between objects and the scene. Comprehensive experiments demonstrate our method's superiority in generating physically plausible, text-compliant interactions compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。