arXiv:2410.19586cs.MMcs.CV2024-10IJCV被引 3

让手语翻译生成多种合理文本,提升自然度与多样性。

Diverse Sign Language Translation

  • 用大模型生成多参考文本,仅由母语者校正,大幅提升标注效率。
  • 引入最大奖励强化学习,使翻译结果在多样性和准确性上双优。
  • 为手语翻译新任务提供基准模型与评估指标,适合研究多答案生成的学者。

与口语类似,单一手语表达可对应多个有效文本解释。因此,学习刚性的一对一映射可能不足,尤其在数据有限时。本文提出多样手语翻译(DivSLT)任务,旨在生成多样且准确的文本翻译。首先,利用大语言模型(LLM)为广泛使用的CSL-Daily和PHOENIX14T数据集生成多个参考译文,仅由母语者修正错误,显著提升标注效率。其次,提出基准模型以推动该任务研究:探索多参考训练策略以实现多样化翻译;为提升准确性,采用最大奖励驱动的强化学习目标,最大化翻译结果的奖励。同时,设计多种评估指标衡量准确性、多样性和语义精确度。在增强数据集上的实验表明,该方法不仅翻译性能更优,且生成结果更具多样性。

原文摘要 · Abstract (English)

Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language translation (SLT) models might be inadequate, particularly in the case of limited data. In this work, we introduce a Diverse Sign Language Translation (DivSLT) task, aiming to generate diverse yet accurate translations for sign language videos. Firstly, we employ large language models (LLM) to generate multiple references for the widely-used CSL-Daily and PHOENIX14T SLT datasets. Here, native speakers are only invited to touch up inaccurate references, thus significantly improving the annotation efficiency. Secondly, we provide a benchmark model to spur research in this task. Specifically, we investigate multi-reference training strategies to enable our DivSLT model to achieve diverse translations. Then, to enhance translation accuracy, we employ the max-reward-driven reinforcement learning objective that maximizes the reward of the translated result. Additionally, we utilize multiple metrics to assess the accuracy, diversity, and semantic precision of the DivSLT task. Experimental results on the enriched datasets demonstrate that our DivSLT method achieves not only better translation performance but also diverse translation results.

手语翻译多答案生成大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。