arXiv:2501.07295cs.RO2025-01被引 15

用大模型让机器人理解各种手势,包括影视经典动作。

GestLLM: Advanced Hand Gesture Interpretation via Large Language Models for Human-Robot Interaction

  • 结合MediaPipe与大语言模型,解析多样手势
  • 无需预训练即可识别《星际迷航》等特殊手势
  • 适合人机协作、辅助机器人与互动娱乐场景

本文提出GestLLM,一种通过手部手势实现人机交互的先进系统。与依赖预设手势的传统方法不同,GestLLM利用MediaPipe进行特征提取,并结合大语言模型,可理解多种多样的手势,包括非标准或流行文化中的手势(如《星际迷航》的瓦肯致意)。该系统在不需额外预训练或提示工程的情况下,实现了与领先视觉-语言模型相当的性能,支持传统数据集未充分覆盖的手势。其灵活性显著提升了人机交互的自然性与包容性,为高级人机协作、辅助机器人及互动娱乐提供了有效解决方案。

原文摘要 · Abstract (English)

This paper introduces GestLLM, an advanced system for human-robot interaction that enables intuitive robot control through hand gestures. Unlike conventional systems, which rely on a limited set of predefined gestures, GestLLM leverages large language models and feature extraction via MediaPipe to interpret a diverse range of gestures. This integration addresses key limitations in existing systems, such as restricted gesture flexibility and the inability to recognize complex or unconventional gestures commonly used in human communication. By combining state-of-the-art feature extraction and language model capabilities, GestLLM achieves performance comparable to leading vision-language models while supporting gestures underrepresented in traditional datasets. For example, this includes gestures from popular culture, such as the ``Vulcan salute" from Star Trek, without any additional pretraining, prompt engineering, etc. This flexibility enhances the naturalness and inclusivity of robot control, making interactions more intuitive and user-friendly. GestLLM provides a significant step forward in gesture-based interaction, enabling robots to understand and respond to a wide variety of hand gestures effectively. This paper outlines its design, implementation, and evaluation, demonstrating its potential applications in advanced human-robot collaboration, assistive robotics, and interactive entertainment.

手势识别人机交互大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。