arXiv:2506.20795cs.CVcs.HC2025-06被引 4

对比视觉大模型与骨骼模型在人机交互手势识别中的表现

How do Foundation Models Compare to Skeleton-Based Approaches for Gesture Recognition in Human-Robot Interaction?

  • 用视觉大模型和骨骼模型对比手势识别效果
  • 骨骼模型最优,视觉大模型仅差一点且更简单
  • 适合想简化系统的人机交互研究者

手势使非语言人机通信成为可能,尤其在嘈杂的敏捷生产环境中。传统深度学习方法依赖特定任务架构,以图像、视频或骨骼姿态为输入。而具备强泛化能力的视觉基础模型(VFMs)和视觉语言模型(VLMs)有望通过替代专用模块来降低系统复杂度。本研究评估了动态全身手势识别中 V-JEPA(先进视觉基础模型)、Gemini Flash 2.0(多模态视觉语言模型)与 HD-GCN(顶尖骨骼模型)的表现。我们引入 NUGGET 数据集,专门用于评估人机物流场景中的手势识别。实验表明,HD-GCN 性能最佳,但 V-JEPA 仅需简单分类头即可接近其表现,展示了作为共享多任务模型减少系统复杂性的可能;相比之下,Gemini 在零样本设置下仅靠文本描述难以区分手势,凸显手势输入表示仍需深入研究。

原文摘要 · Abstract (English)

Gestures enable non-verbal human-robot communication, especially in noisy environments like agile production. Traditional deep learning-based gesture recognition relies on task-specific architectures using images, videos, or skeletal pose estimates as input. Meanwhile, Vision Foundation Models (VFMs) and Vision Language Models (VLMs) with their strong generalization abilities offer potential to reduce system complexity by replacing dedicated task-specific modules. This study investigates adapting such models for dynamic, full-body gesture recognition, comparing V-JEPA (a state-of-the-art VFM), Gemini Flash 2.0 (a multimodal VLM), and HD-GCN (a top-performing skeleton-based approach). We introduce NUGGET, a dataset tailored for human-robot communication in intralogistics environments, to evaluate the different gesture recognition approaches. In our experiments, HD-GCN achieves best performance, but V-JEPA comes close with a simple, task-specific classification head - thus paving a possible way towards reducing system complexity, by using it as a shared multi-task model. In contrast, Gemini struggles to differentiate gestures based solely on textual descriptions in the zero-shot setting, highlighting the need of further research on suitable input representations for gestures.

手势识别视觉大模型人机交互骨骼模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。