arXiv:2504.09623cs.CVcs.AI2025-04CVPR被引 13

让AI理解人指物+语言描述,提升3D场景定位准确率

Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

  • 用数据增强生成带手势的3D语言定位数据
  • 新模型在3D定位任务上提升30%准确率
  • 适合研究具身智能与人机交互的学者

三维具身参考理解(3D-ERU)结合语言描述和人类指向手势,以定位3D场景中目标物体。尽管已有纯语言基3D定位工作,但对融合手势的3D-ERU研究仍有限。为此,我们提出数据增强框架Imputer,将其应用于仅含语言指令的现有3D场景数据集,构建了新基准数据集ImputeRefer,引入人类指向手势。同时,我们提出Ges3ViG模型,在3D-ERU任务中相比其他模型提升约30%准确率,相较于纯语言基模型提升约9%。代码与数据集已公开于https://github.com/AtharvMane/Ges3ViG。

原文摘要 · Abstract (English)

3-Dimensional Embodied Reference Understanding (3D-ERU) combines a language description and an accompanying pointing gesture to identify the most relevant target object in a 3D scene. Although prior work has explored pure language-based 3D grounding, there has been limited exploration of 3D-ERU, which also incorporates human pointing gestures. To address this gap, we introduce a data augmentation framework-Imputer, and use it to curate a new benchmark dataset-ImputeRefer for 3D-ERU, by incorporating human pointing gestures into existing 3D scene datasets that only contain language instructions. We also propose Ges3ViG, a novel model for 3D-ERU that achieves ~30% improvement in accuracy as compared to other 3D-ERU models and ~9% compared to other purely language-based 3D grounding models. Our code and dataset are available at https://github.com/AtharvMane/Ges3ViG.

3D定位具身智能手势理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。