arXiv:2509.10266cs.CVcs.AI2025-09被引 2

融合口型信息提升手语翻译准确率

SignMouth: Leveraging Mouthing Cues for Sign Language Translation by Multimodal Contrastive Fusion

  • 通过空间手势与唇动特征融合提升识别精度
  • 在PHOENIX14T数据集上BLEU-4达24.71,ROUGE达48.38
  • 适合关注非手动信号的多模态研究者

手语翻译旨在将手语视频中的自然语言转化为文本,是实现包容性沟通的重要桥梁。尽管近年来利用强大视觉主干网络和大语言模型取得进展,但多数方法仍主要关注手部动作,忽视了口型等非手动线索。事实上,口型在手语中传递关键语言信息,对区分视觉相似的手势至关重要。本文提出SignClip框架,融合手部动作与口型运动特征,并引入分层对比学习机制,实现手语-唇动与视觉-文本模态间的语义一致性对齐。在两个基准数据集PHOENIX14T和How2Sign上的实验表明,该方法具有显著优势。例如,在PHOENIX14T的无词典设置下,相比先前最优模型SpaMo,BLEU-4从24.32提升至24.71,ROUGE从46.57提升至48.38。

原文摘要 · Abstract (English)

Sign language translation (SLT) aims to translate natural language from sign language videos, serving as a vital bridge for inclusive communication. While recent advances leverage powerful visual backbones and large language models, most approaches mainly focus on manual signals (hand gestures) and tend to overlook non-manual cues like mouthing. In fact, mouthing conveys essential linguistic information in sign languages and plays a crucial role in disambiguating visually similar signs. In this paper, we propose SignClip, a novel framework to improve the accuracy of sign language translation. It fuses manual and non-manual cues, specifically spatial gesture and lip movement features. Besides, SignClip introduces a hierarchical contrastive learning framework with multi-level alignment objectives, ensuring semantic consistency across sign-lip and visual-text modalities. Extensive experiments on two benchmark datasets, PHOENIX14T and How2Sign, demonstrate the superiority of our approach. For example, on PHOENIX14T, in the Gloss-free setting, SignClip surpasses the previous state-of-the-art model SpaMo, improving BLEU-4 from 24.32 to 24.71, and ROUGE from 46.57 to 48.38.

手语翻译多模态融合口型识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。