arXiv:2510.21814cs.CVcs.AI2025-10被引 2

用大模型实现实时自由手势理解,准确率和速度远超现有方案

Gestura: A LVLM-Powered System Bridging Motion and Semantics for Real-Time Free-Form Gesture Understanding

  • 基于大视觉语言模型,融合手部解剖先验与思维链推理
  • 支持30万+带问答标注的自由手势数据,识别准确率显著提升
  • 适合交互设计、人机协同等需自然手势输入的场景

自由手势理解在人机交互中极具吸引力,因其摆脱了预定义手势类别的限制。然而,现有唯一解决方案GestureGPT存在识别准确率低、响应慢的问题。本文提出Gestura,一个端到端的自由手势理解系统。Gestura利用预训练的大视觉语言模型(LVLM)将高度动态多样的自由手势模式与高层语义概念对齐。为更好捕捉不同风格下的细微手部动作,引入关键点处理模块,通过嵌入解剖学手部先验知识弥补LVLM在细粒度领域知识上的不足。进一步采用思维链(CoT)推理策略,实现逐步语义推断,将浅层知识转化为深层语义理解,显著提升模型对模糊或非常规手势的解析能力。上述组件协同工作,使Gestura具备鲁棒且可适应的自由手势理解能力。此外,我们构建了首个面向自由手势意图推理与理解的开源数据集,包含超过30万条标注的问答对。

原文摘要 · Abstract (English)

Free-form gesture understanding is highly appealing for human-computer interaction, as it liberates users from the constraints of predefined gesture categories. However, the sole existing solution GestureGPT suffers from limited recognition accuracy and slow response times. In this paper, we propose Gestura, an end-to-end system for free-form gesture understanding. Gestura harnesses a pre-trained Large Vision-Language Model (LVLM) to align the highly dynamic and diverse patterns of free-form gestures with high-level semantic concepts. To better capture subtle hand movements across different styles, we introduce a Landmark Processing Module that compensate for LVLMs' lack of fine-grained domain knowledge by embedding anatomical hand priors. Further, a Chain-of-Thought (CoT) reasoning strategy enables step-by-step semantic inference, transforming shallow knowledge into deep semantic understanding and significantly enhancing the model's ability to interpret ambiguous or unconventional gestures. Together, these components allow Gestura to achieve robust and adaptable free-form gesture comprehension. Additionally, we have developed the first open-source dataset for free-form gesture intention reasoning and understanding with over 300,000 annotated QA pairs.

手势识别大模型人机交互视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。