arXiv:2503.02421cs.LG2025-03被引 1

首个面向希腊手语生成的基于Transformer的框架,实现文本与手语动作互译。

A Transformer-Based Framework for Greek Sign Language Production using Extended Skeletal Motion Representations

  • 用Transformer模型将文本转为人体骨骼关键点序列。
  • 在Elementary23数据集上实现高质量手语视频生成,效果优于基线。
  • 适合研究手语合成、无障碍沟通系统开发的团队参考。

手语是全球聋人社区的主要交流方式。为打破聋人/重听者与听力人群之间的沟通障碍,亟需构建能实现口语与手语相互转换的系统。基于前期研究洞察,我们提出首个面向希腊手语生成(Greek SLP)的深度学习模型。该模型采用基于Transformer的架构,支持从文本输入生成手语动作关键点序列,反之亦然。我们在希腊手语数据集Elementary23上进行评估,通过一系列对比分析与消融实验验证了该流水线的有效性。其核心组件——数据驱动的词素生成、视频到文本的训练策略,以及教师强制-自回归解码调度算法,显著提升了生成手语视频的质量。

原文摘要 · Abstract (English)

Sign Languages are the primary form of communication for Deaf communities across the world. To break the communication barriers between the Deaf and Hard-of-Hearing and the hearing communities, it is imperative to build systems capable of translating the spoken language into sign language and vice versa. Building on insights from previous research, we propose a deep learning model for Sign Language Production (SLP), which to our knowledge is the first attempt on Greek SLP. We tackle this task by utilizing a transformer-based architecture that enables the translation from text input to human pose keypoints, and the opposite. We evaluate the effectiveness of the proposed pipeline on the Greek SL dataset Elementary23, through a series of comparative analyses and ablation studies. Our pipeline's components, which include data-driven gloss generation, training through video to text translation and a scheduling algorithm for teacher forcing - auto-regressive decoding seem to actively enhance the quality of produced SL videos.

手语生成Transformer骨骼表示无障碍

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。