用大模型生成手语翻译的多种参考句,提升评估准确性。
Beyond a Single Reference: Training and Evaluation with Paraphrases in Sign Language Translation
- 用大模型自动生成手语翻译的多种语义等价表述作为参考。
- 评估时引入多参考句,使自动指标与人工评分更一致。
- 提出BLEUpara新指标,适合评估手语翻译系统质量。
多数手语翻译数据集将每个手语表达对应单一书面语参考句,但手语与口语间存在高度非同构性,同一手语可有多个合理翻译。这限制了模型训练与评估,尤其影响基于n-gram的BLEU等指标。本文研究利用大语言模型自动生成书面语翻译的改写版本,作为手语翻译的合成替代参考句。首先,通过改进的ParaScore评估多种改写策略与模型效果;其次,在YouTubeASL和How2Sign数据集上,研究改写句对基于姿态的T5模型训练与评估的影响。结果表明,直接在训练中使用改写句无法提升性能,甚至有害;但在评估阶段引入改写句,能显著提高自动评分,并更贴近人工判断。为此,提出BLEUpara,一种基于多改写参考句的扩展版BLEU指标。人工评估验证,BLEUpara与翻译质量感知相关性更强。论文公开所有生成的改写句、生成与评估代码,支持可复现、更可靠的SLT系统评估。
原文摘要 · Abstract (English)
Most Sign Language Translation (SLT) corpora pair each signed utterance with a single written-language reference, despite the highly non-isomorphic relationship between sign and spoken languages, where multiple translations can be equally valid. This limitation constrains both model training and evaluation, particularly for n-gram-based metrics such as BLEU. In this work, we investigate the use of Large Language Models to automatically generate paraphrased variants of written-language translations as synthetic alternative references for SLT. First, we compare multiple paraphrasing strategies and models using an adapted ParaScore metric. Second, we study the impact of paraphrases on both training and evaluation of the pose-based T5 model on the YouTubeASL and How2Sign datasets. Our results show that naively incorporating paraphrases during training does not improve translation performance and can even be detrimental. In contrast, using paraphrases during evaluation leads to higher automatic scores and better alignment with human judgments. To formalize this observation, we introduce BLEUpara, an extension of BLEU that evaluates translations against multiple paraphrased references. Human evaluation confirms that BLEUpara correlates more strongly with perceived translation quality. We release all generated paraphrases, generation and evaluation code to support reproducible and more reliable evaluation of SLT systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。