通过分层对齐视频与伪词义,提升无词义标注的手语翻译质量。
Hierarchical Feature Alignment for Gloss-Free Sign Language Translation
- 分三层提取帧、片段和视频级特征,与伪词义对齐
- 在CVL和PHOENIX14T数据集上,BLEU-4提升1.8~3.2点
- 适合无词义标注场景,兼顾精度与推理效率
手语翻译(SLT)旨在将手语视频转换为口语句子。现有方法在端到端学习中常因视觉与文本表征差异而表现不佳。基于词义的方法虽能缓解该问题,但依赖人工标注;无词义方法更灵活,但需有效对齐策略。近期大语言模型使无词义SLT成为可能,可从手语视频生成类文本表示。本文提出一种受手语结构启发的分层预训练策略,引入伪词义并结合对比视频-语言对齐。方法在帧、片段和视频三个层级上提取特征,并分别与伪词义及目标语句对齐,从而提升翻译质量。实验表明,本方法在保持高效的同时,显著提升BLEU-4与ROUGE分数。
原文摘要 · Abstract (English)
Sign Language Translation (SLT) attempts to convert sign language videos into spoken sentences. However, many existing methods struggle with the disparity between visual and textual representations during end-to-end learning. Gloss-based approaches help to bridge this gap by leveraging structured linguistic information. While, gloss-free methods offer greater flexibility and remove the burden of annotation, they require effective alignment strategies. Recent advances in Large Language Models (LLMs) have enabled gloss-free SLT by generating text-like representations from sign videos. In this work, we introduce a novel hierarchical pre-training strategy inspired by the structure of sign language, incorporating pseudo-glosses and contrastive video-language alignment. Our method hierarchically extracts features at frame, segment, and video levels, aligning them with pseudo-glosses and the spoken sentence to enhance translation quality. Experiments demonstrate that our approach improves BLEU-4 and ROUGE scores while maintaining efficiency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。