arXiv:2411.16789cs.CVcs.CL2024-11ICCV被引 16

用现成大模型直接翻译手语,无需中间标注词

Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation

  • 用多模态大模型生成手语动作的详细描述
  • 在PHOENIX14T和CSL-Daily上达到最新最好效果
  • 适合做无障碍沟通、手语技术研究者

手语翻译(SLT)是将手语图像转换为口语语言的挑战性任务。为成功完成此任务,模型需弥合模态差距,并识别手语组件中的细微差异以准确理解其含义。为此,我们提出一种名为MMSLT的新颖无标注词手语翻译框架,利用现成多模态大语言模型(MLLMs)的表征能力。具体地,我们使用MLLMs生成手语组件的详细文本描述,再通过提出的多模态-语言预训练模块,将这些描述特征与手语视频特征融合,使其在口语句子空间中对齐。该方法在基准数据集PHOENIX14T和CSL-Daily上实现最先进性能,凸显了MLLM在手语翻译中的有效潜力。代码已开源。

原文摘要 · Abstract (English)

Sign language translation (SLT) is a challenging task that involves translating sign language images into spoken language. For SLT models to perform this task successfully, they must bridge the modality gap and identify subtle variations in sign language components to understand their meanings accurately. To address these challenges, we propose a novel gloss-free SLT framework called Multimodal Sign Language Translation (MMSLT), which leverages the representational capabilities of off-the-shelf multimodal large language models (MLLMs). Specifically, we use MLLMs to generate detailed textual descriptions of sign language components. Then, through our proposed multimodal-language pre-training module, we integrate these description features with sign video features to align them within the spoken sentence space. Our approach achieves state-of-the-art performance on benchmark datasets PHOENIX14T and CSL-Daily, highlighting the potential of MLLMs to be utilized effectively in SLT. Code is available at https://github.com/hwjeon98/MMSLT.

手语翻译多模态大模型无标注词视觉语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。