用语言模板生成假数据,让手势翻译模型更准
POSESTITCH-SLT: Linguistically Inspired Pose-Stitching for End-to-End Sign Language Translation
- 用语言模板生成句子对,给模型提供合成监督信号
- 在两个数据集上翻译准确率提升超2倍,最高达3.43
- 适合资源少的手势语翻译研究者参考
手势语翻译因缺乏大规模句对齐数据而困难。现有方法多聚焦特征提取与结构改进。本文提出基于语言模板的预训练方法POSESTITCH-SLT,利用模板生成句子对进行训练。在How2Sign和iSign两个数据集上,基于Transformer的简单编码器-解码器架构取得显著提升:BLEU-4得分从1.97增至4.56(How2Sign),从0.55增至3.43(iSign),超越现有基于姿态的无词元翻译方法。结果证明,在低资源场景下,模板驱动的合成监督有效提升了模型性能。
原文摘要 · Abstract (English)
Sign language translation remains a challenging task due to the scarcity of large-scale, sentence-aligned datasets. Prior arts have focused on various feature extraction and architectural changes to support neural machine translation for sign languages. We propose POSESTITCH-SLT, a novel pre-training scheme that is inspired by linguistic-templates-based sentence generation technique. With translation comparison on two sign language datasets, How2Sign and iSign, we show that a simple transformer-based encoder-decoder architecture outperforms the prior art when considering template-generated sentence pairs in training. We achieve BLEU-4 score improvements from 1.97 to 4.56 on How2Sign and from 0.55 to 3.43 on iSign, surpassing prior state-of-the-art methods for pose-based gloss-free translation. The results demonstrate the effectiveness of template-driven synthetic supervision in low-resource sign language settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。