用大模型直接翻译3D手语,提升聋人沟通无障碍体验
Large Sign Language Models: Toward 3D American Sign Language Translation
- 以大语言模型为基底,直接处理3D手语数据,捕捉空间与深度信息
- 支持直接翻译与指令引导翻译,灵活应对不同表达需求
- 推动多模态语言融入大模型,迈向更包容的智能系统
我们提出大型手语模型(LSLM),一种基于大语言模型(LLM)框架的3D美国手语(ASL)翻译方法,旨在改善听力障碍者在虚拟交流中的体验。与依赖2D视频的现有手语识别方法不同,本方法直接利用3D手语数据,捕捉3D场景中丰富的空间、手势和深度信息,从而实现更精准且鲁棒的翻译,提升数字通信对听障群体的可及性。此外,本工作探索将复杂、具身的多模态语言整合进大语言模型的处理能力中,突破纯文本输入的局限,拓展其对人类沟通的理解。研究涵盖从3D动作特征直接翻译成文本,以及通过外部提示调节翻译的指令引导设置,提供更高灵活性。该工作为构建能理解多样语言形式的包容性多模态智能系统奠定了基础。
原文摘要 · Abstract (English)
We present Large Sign Language Models (LSLM), a novel framework for translating 3D American Sign Language (ASL) by leveraging Large Language Models (LLMs) as the backbone, which can benefit hearing-impaired individuals' virtual communication. Unlike existing sign language recognition methods that rely on 2D video, our approach directly utilizes 3D sign language data to capture rich spatial, gestural, and depth information in 3D scenes. This enables more accurate and resilient translation, enhancing digital communication accessibility for the hearing-impaired community. Beyond the task of ASL translation, our work explores the integration of complex, embodied multimodal languages into the processing capabilities of LLMs, moving beyond purely text-based inputs to broaden their understanding of human communication. We investigate both direct translation from 3D gesture features to text and an instruction-guided setting where translations can be modulated by external prompts, offering greater flexibility. This work provides a foundational step toward inclusive, multimodal intelligent systems capable of understanding diverse forms of language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。