arXiv:2507.04465cs.CV2025-07综述被引 13

系统梳理手势识别的主流方法、数据集与未来方向,助研究者快速掌握领域全貌。

Visual Hand Gesture Recognition with Deep Learning: A Comprehensive Review of Methods, Datasets, Challenges and Future Research Directions

  • 按静态、孤立动态、连续手势三类任务分类整理现有方法
  • 归纳当前最优模型的架构趋势与学习策略
  • 总结领域核心挑战,指明未来研究可行路径

深度学习模型的快速发展和大规模数据集的涌现,推动了视觉手部手势识别(VHGR)领域的持续关注,并催生了手语理解、人机交互等广泛应用。尽管已有大量研究成果,但针对VHGR的系统性综述仍属空白,研究人员需在数百篇论文中自行梳理最新进展。本综述旨在填补这一空白,通过系统的文献调研与结构化呈现,全面概述该计算机视觉领域。重点回答四个关键问题:主要研究维度是什么?当前最先进方法有哪些?不同方法与任务间有何对比洞察?未来研究面临哪些挑战?综述基于分类体系组织关键技术,将最先进方法分为静态、孤立动态和连续手势识别三类,分别梳理其架构演进与学习策略。同时回顾常用数据集与标准评估指标,为未来方法提供实验基准。最后,指出包括通用计算机视觉问题与领域特有难题在内的主要挑战,并提出具有前景的研究方向。

原文摘要 · Abstract (English)

The rapid evolution of deep learning (DL) models and the ever-increasing size of available datasets have raised the interest of the research community in the always-important field of visual hand gesture recognition (VHGR), and delivered a wide range of applications, such as sign language understanding and human-computer interaction. Despite the large volume of research works in the field, a structured and complete survey on VHGR is still missing, leaving researchers to navigate through hundreds of papers in order to find the current state-of-the-art (SOTA). The current survey aims to fill this gap by presenting a comprehensive overview of this computer vision field. With a systematic research methodology and a structured presentation of the various methods, datasets, and evaluation metrics, this review aims to constitute a useful guideline for researchers, helping them to propose improvements. Specifically, this survey focuses on four fundamental questions: what are the main VHGR aspects, what are the current SOTA methods, what comparative insights can be drawn across methods and tasks, and which challenges shape future research. Starting with the methodology used to locate the related literature, the survey identifies and organizes the key VHGR approaches in a taxonomy-based format. The SOTA methods are grouped across three primary VHGR tasks: static, isolated dynamic and continuous gesture recognition. For each task, the architectural trends and learning strategies are listed. To support the experimental evaluation of future methods in the field, the study reviews commonly used datasets and presents the standard performance metrics. Our survey concludes by identifying the major challenges in VHGR, including both general computer vision issues and domain-specific obstacles, and outlines promising directions for future research.

手势识别综述深度学习人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。