无需手语标注,直接将手势转为机器人指令,实现无障碍实时操控。
SignVLA: A Gloss-Free Vision-Language-Action Framework for Real-Time Sign Language-Guided Robotic Manipulation
- 跳过中间标注环节,直接从手势图像生成语义指令。
- 实测在多种场景下能精准驱动机器人执行动作。
- 适合残障人士或安全关键环境中的无障碍人机交互。
我们提出首个无需手语标注的视觉-语言-动作(VLA)框架,实现直观且包容的人机交互。与依赖词素标注的传统方法不同,该系统采用无词素范式,直接将视觉手语手势映射为语义指令,降低标注成本并避免词素表示带来的信息损失,提升多模态交互的自然性与可扩展性。本文聚焦于实时字母级指拼接口,为机器人控制提供鲁棒、低延迟的通信通道。相较于大规模连续手语识别,字母级交互在安全性要求高的具身环境中具有更高的可靠性、可解释性与部署可行性。所提流程通过几何归一化、时间平滑和词汇优化,将连续手势流转化为连贯的语言命令,确保交互稳定一致。此外,该框架支持未来集成基于Transformer的无词素手语模型,实现词级与句级语义理解。实验结果表明,在多种交互场景下,该系统能有效将手语指令精准落地为机器人动作,展现了其在可访问、可扩展、多模态具身智能方面的潜力。
原文摘要 · Abstract (English)
We present, to our knowledge, the first sign language-driven Vision-Language-Action (VLA) framework for intuitive and inclusive human-robot interaction. Unlike conventional approaches that rely on gloss annotations as intermediate supervision, the proposed system adopts a gloss-free paradigm and directly maps visual sign gestures to semantic instructions. This design reduces annotation cost and avoids the information loss introduced by gloss representations, enabling more natural and scalable multimodal interaction. In this work, we focus on a real-time alphabet-level finger-spelling interface that provides a robust and low-latency communication channel for robotic control. Compared with large-scale continuous sign language recognition, alphabet-level interaction offers improved reliability, interpretability, and deployment feasibility in safety-critical embodied environments. The proposed pipeline transforms continuous gesture streams into coherent language commands through geometric normalization, temporal smoothing, and lexical refinement, ensuring stable and consistent interaction. Furthermore, the framework is designed to support future integration of transformer-based gloss-free sign language models, enabling scalable word-level and sentence-level semantic understanding. Experimental results demonstrate the effectiveness of the proposed system in grounding sign-derived instructions into precise robotic actions under diverse interaction scenarios. These results highlight the potential of the framework to advance accessible, scalable, and multimodal embodied intelligence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。