从一张图和一句话生成逼真的3D手物交互轨迹。
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
- 用扩散模型结合图像与文本,检索相似3D物体并约束交互几何。
- 在HOI4D和H2O数据集上生成轨迹的多样性和物理合理性更优。
- 适合做机器人操作、具身智能的研究者参考。
当人类抓取物体时,会在脑海中形成操作轨迹。建模手物交互先验对推动机器人与具身AI系统在真实世界中的有效运作具有重要意义。我们提出SIGHT任务,即从单张图像和简短语言描述中生成真实且物理合理的3D手物交互轨迹。以往工作通常依赖缺乏目标物体显式关联的文本输入,或假设可访问3D物体网格,而后者获取难度远高于2D图像。我们提出SIGHT-Fusion,一种基于扩散的图像-文本条件生成模型,通过从数据库中检索最相似的3D物体网格,并在推理阶段引入新颖的几何引导机制,强制实现手物交互约束。我们在HOI4D和H2O数据集上评估模型,并适配相关基线方法。实验表明,我们的模型在生成轨迹的多样性、质量以及手物交互几何指标上均表现更优。
原文摘要 · Abstract (English)
When humans grasp an object, they naturally form trajectories in their minds to manipulate it for specific tasks. Modeling hand-object interaction priors holds significant potential to advance robotic and embodied AI systems in learning to operate effectively within the physical world. We introduce SIGHT, a novel task focused on generating realistic and physically plausible 3D hand-object interaction trajectories from a single image and a brief language-based task description. Prior work on hand-object trajectory generation typically relies on textual input that lacks explicit grounding to the target object, or assumes access to 3D object meshes, which are often considerably more difficult to obtain than 2D images. We propose SIGHT-Fusion, a novel diffusion-based image-text conditioned generative model that tackles this task by retrieving the most similar 3D object mesh from a database and enforcing geometric hand-object interaction constraints via a novel inference-time diffusion guidance. We benchmark our model on the HOI4D and H2O datasets, adapting relevant baselines for this novel task. Experiments demonstrate our superior performance in the diversity and quality of generated trajectories, as well as in hand-object interaction geometry metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。