arXiv:2506.09643cs.CLcs.CV2025-06被引 8

用手语生成数据增强手语翻译模型性能,提升最多19%。

Using Sign Language Production as Data Augmentation to enhance Sign Language Translation

  • 用骨骼驱动和真实感生成模型合成手语视频
  • 在多个数据集上使翻译准确率提升最高达19%
  • 适合资源有限但需改进手语翻译的团队

机器学习模型依赖大量高质量数据,但手语因成本高、数据少且涉及隐私,属于低资源语言,现有数据集规模远小于口语数据。本文提出利用手语生成技术增强已有手语数据集,以提升手语翻译模型性能。采用三种方法:基于骨骼的手语生成、手语拼接,以及两种逼真生成模型(SignGAN 和 SignSplat)。通过生成不同签员外貌与骨骼动作变化,扩展数据多样性。实验表明,该方法可有效提升手语翻译模型性能,最高提升19%,为资源受限环境下的手语翻译系统提供更鲁棒、准确的解决方案。

原文摘要 · Abstract (English)

Machine learning models fundamentally rely on large quantities of high-quality data. Collecting the necessary data for these models can be challenging due to cost, scarcity, and privacy restrictions. Signed languages are visual languages used by the deaf community and are considered low-resource languages. Sign language datasets are often orders of magnitude smaller than their spoken language counterparts. Sign Language Production is the task of generating sign language videos from spoken language sentences, while Sign Language Translation is the reverse translation task. Here, we propose leveraging recent advancements in Sign Language Production to augment existing sign language datasets and enhance the performance of Sign Language Translation models. For this, we utilize three techniques: a skeleton-based approach to production, sign stitching, and two photo-realistic generative models, SignGAN and SignSplat. We evaluate the effectiveness of these techniques in enhancing the performance of Sign Language Translation models by generating variation in the signer's appearance and the motion of the skeletal data. Our results demonstrate that the proposed methods can effectively augment existing datasets and enhance the performance of Sign Language Translation models by up to 19%, paving the way for more robust and accurate Sign Language Translation systems, even in resource-constrained environments.

手语翻译数据增强生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。