用短片段拼接生成手语语料,低成本提升翻译模型性能
Less is more: concatenating videos for Sign Language Translation from a small set of signs
- 通过拼接孤立手语片段生成大规模语料库
- 在50万视频数据下达到BLEU-4 9.2%、METEOR 26.2%
- 适合资源受限的手语翻译研究者参考
由于视频采集和标注成本高昂,训练巴西手语(Libras)到葡萄牙语翻译模型面临标注数据稀缺的挑战。本文提出通过拼接包含孤立手势的短片段来生成手语内容,用于训练翻译模型。我们使用V-LIBRASIL数据集(含4,089个手语视频,覆盖1,364个手势,由至少三人解读),生成了约17万、30万和50万条句子及其对应的Libras翻译,用于模型训练。实验结果表明,在该数据规模下,模型在BLEU-4上取得9.2%的得分,在METEOR上达到26.2%。该方法显著降低了数据构建成本,为未来手语翻译研究提供了可行路径。
原文摘要 · Abstract (English)
The limited amount of labeled data for training the Brazilian Sign Language (Libras) to Portuguese Translation models is a challenging problem due to video collection and annotation costs. This paper proposes generating sign language content by concatenating short clips containing isolated signals for training Sign Language Translation models. We employ the V-LIBRASIL dataset, composed of 4,089 sign videos for 1,364 signs, interpreted by at least three persons, to create hundreds of thousands of sentences with their respective Libras translation, and then, to feed the model. More specifically, we propose several experiments varying the vocabulary size and sentence structure, generating datasets with approximately 170K, 300K, and 500K videos. Our results achieve meaningful scores of 9.2% and 26.2% for BLEU-4 and METEOR, respectively. Our technique enables the creation or extension of existing datasets at a much lower cost than the collection and annotation of thousands of sentences providing clear directions for future works.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。