arXiv:2603.15403cs.CV2026-03

用单图识别手指指向的物体,提升人机交互自然性

Pointing-Based Object Recognition

  • 融合姿态、深度与视觉语言模型,从单图定位指向目标
  • 加入深度信息后,复杂场景下识别准确率显著提升
  • 模块化设计,无需专用深度传感器也能部署

本文提出一个完整的管道,用于通过RGB图像识别由人类指向手势所指示的物体。随着人机交互向更直观的界面发展,识别非语言交流的目标变得至关重要。所提系统整合了多项现有最先进方法,包括目标检测、人体姿态估计、单目深度估计和视觉-语言模型。我们评估了从单张图像重建的三维空间信息以及图像描述模型在纠正分类错误中的作用。在自建数据集上的实验结果表明,引入深度信息能显著提升目标识别效果,尤其是在存在重叠物体的复杂场景中。该方法的模块化设计使其可在缺乏专用深度传感器的环境中部署。

原文摘要 · Abstract (English)

This paper presents a comprehensive pipeline for recognizing objects targeted by human pointing gestures using RGB images. As human-robot interaction moves toward more intuitive interfaces, the ability to identify targets of non-verbal communication becomes crucial. Our proposed system integrates several existing state-of-the-art methods, including object detection, body pose estimation, monocular depth estimation, and vision-language models. We evaluate the impact of 3D spatial information reconstructed from a single image and the utility of image captioning models in correcting classification errors. Experimental results on a custom dataset show that incorporating depth information significantly improves target identification, especially in complex scenes with overlapping objects. The modularity of the approach allows for deployment in environments where specialized depth sensors are unavailable.

目标识别人机交互单目深度视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。