用手语实时指挥机器人,让听障用户也能无障碍操控。
SignVLA: Real-Time Sign Language-Guided Robotic Manipulation via Attention LSTM and Vision-Language-Action Models

- 手语转文本:用注意力LSTM捕捉手势时序动态,识别字母和指令
- 实时稳定识别:在视频流中实现连续手语输入的准确解析
- 无障碍交互:为听障用户设计轻量级手语控制接口,适配智能机器人
视觉-语言-动作(VLA)模型使机器人能根据自然语言指令执行操作任务,但现有接口主要依赖语音或文字输入,对聋哑及言语障碍用户不友好。本文提出SignVLA,一种实时手语引导的VLA框架,支持无障碍人机交互。系统采用模块化手语转文本接口,将视觉手语手势转化为可被下游VLA策略使用的语义指令。基于视频流,SignVLA提取手部关键点特征,利用增强注意力的长短期记忆网络(Attention LSTM)捕捉时序手势动态,实现字母级与命令级手语识别。同时引入时序稳定性模块,提升实时交互中的预测一致性。生成的指令序列随后输入下游VLA策略,完成手语驱动的机器人操作。实验表明,系统具备稳定的实时手语识别能力,并成功执行各类操作任务。研究结果表明,轻量级时序手语识别可作为多模态具身智能的有效且实用的无障碍接入层。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models enable robots to execute manipulation tasks from natural-language instructions grounded in visual observations. However, existing VLA interfaces primarily rely on speech or text input, limiting accessibility for deaf, hard-of-hearing, and speech-impaired users. We present SignVLA, a real-time sign-language-guided VLA framework for accessible human-robot interaction. The system introduces a modular sign-to-text interface that converts visual sign gestures into semantic instructions compatible with downstream VLA policies. Given video streams, SignVLA extracts hand landmark features and employs an attention-enhanced Long Short-Term Memory (LSTM) network to capture temporal gesture dynamics for alphabet- and command-level sign recognition. A temporal stabilization module further improves prediction consistency in real-time interaction settings.The generated instruction sequence is then passed to a downstream VLA policy for sign-conditioned robotic manipulation. Experimental results demonstrate stable real-time sign recognition and successful execution of manipulation tasks driven by sign-language inputs. Our findings suggest that lightweight temporal sign recognition can serve as an effective and practical accessibility layer for multimodal embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。