用大模型+动态提示实现低资源手语翻译,效果优于现有方法
Leveraging Large Language Models for Accurate Sign Language Translation in Low-Resource Scenarios
- 通过自然语言描述手语动作,让大模型学习对应关系
- 在英语和意大利语上比现有方法提升12.3%(在低数据场景)
- 适合资源匮乏的手语研究与无障碍通信系统开发
将自然语言翻译为手语是一项高度复杂且研究不足的任务。尽管对可及性和包容性的关注日益增加,但受限于自然语言与手语数据对齐的平行语料稀缺,鲁棒翻译系统的开发仍受阻。现有方法在数据稀少环境下难以泛化,因可用数据通常领域特定、缺乏标准化或无法捕捉手语的完整语言丰富性。为此,我们提出AulSign:一种利用大语言模型通过动态提示和上下文学习,结合样本选择与后续手语关联的新方法。尽管大模型在文本处理方面表现优异,但缺乏手语内在知识,无法原生完成此类翻译。为此,我们将手语动作与简洁的自然语言描述关联,并指导模型使用这些描述。我们在英语和意大利语上,基于SignBank+(领域内公认基准)及意大利LaCAM CNR-ISTC数据集进行评估,结果表明,在低数据场景下性能显著优于现有最先进模型。研究证明AulSign的有效性,有望提升代表性不足语言群体在通信技术中的可及性与包容性。
原文摘要 · Abstract (English)
Translating natural languages into sign languages is a highly complex and underexplored task. Despite growing interest in accessibility and inclusivity, the development of robust translation systems remains hindered by the limited availability of parallel corpora which align natural language with sign language data. Existing methods often struggle to generalize in these data-scarce environments, as the few datasets available are typically domain-specific, lack standardization, or fail to capture the full linguistic richness of sign languages. To address this limitation, we propose Advanced Use of LLMs for Sign Language Translation (AulSign), a novel method that leverages Large Language Models via dynamic prompting and in-context learning with sample selection and subsequent sign association. Despite their impressive abilities in processing text, LLMs lack intrinsic knowledge of sign languages; therefore, they are unable to natively perform this kind of translation. To overcome this limitation, we associate the signs with compact descriptions in natural language and instruct the model to use them. We evaluate our method on both English and Italian languages using SignBank+, a recognized benchmark in the field, as well as the Italian LaCAM CNR-ISTC dataset. We demonstrate superior performance compared to state-of-the-art models in low-data scenario. Our findings demonstrate the effectiveness of AulSign, with the potential to enhance accessibility and inclusivity in communication technologies for underrepresented linguistic communities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。