arXiv:2507.10972cs.CLcs.CV2025-07中稿 · IEEE ICIP 2025

用分步提示让大模型学会手语生成,提升文本到手语的转换质量。

Teach Me Sign: Stepwise Prompting LLM for Sign Language Production

  • 通过分步提示挖掘大模型中的手语知识,实现文本到手语的准确映射。
  • 在How2Sign和Phoenix14T数据集上显著优于基线方法,手语生成更符合语法规范。
  • 适合关注手语生成、无障碍交互与大模型多模态应用的研究者。

大语言模型凭借强大的推理能力和丰富知识,在多个AI任务中引发变革,但其在手语生成领域的应用仍受限于手语的复杂性与独特规则。本文提出TEAch Me Sign(TEAM-Sign),将手语视为另一种自然语言,通过对大语言模型进行微调,使其学习文本与手语之间的对应关系,从而支持手语生成。考虑到手语与口语的差异,采用分步提示策略,从大模型中提取内在的手语知识,辅助学习与生成过程。在How2Sign和Phoenix14T数据集上的实验结果表明,该方法有效利用了大模型的手语知识与推理能力,能够对齐手语与口语在分布与语法规则上的差异。

原文摘要 · Abstract (English)

Large language models, with their strong reasoning ability and rich knowledge, have brought revolution to many tasks of AI, but their impact on sign language generation remains limited due to its complexity and unique rules. In this paper, we propose TEAch Me Sign (TEAM-Sign), treating sign language as another natural language. By fine-tuning an LLM, we enable it to learn the correspondence between text and sign language, and facilitate generation. Considering the differences between sign and spoken language, we employ a stepwise prompting strategy to extract the inherent sign language knowledge within the LLM, thereby supporting the learning and generation process. Experimental results on How2Sign and Phoenix14T datasets demonstrate that our approach effectively leverages both the sign language knowledge and reasoning capabilities of LLM to align the different distribution and grammatical rules between sign and spoken language.

手语生成大模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。