无需词典标注,用视觉与姿态融合实现手语自动翻译。
ViPo-MLLM: Visual-Pose Multimodal LLM for Gloss-Free Sign Language Translation

- 融合视频与人体姿态信息,通过交叉注意力建模长程依赖。
- 在PHOENIX14T和CSL-Daily数据集上刷新最优结果。
- 适合无词典标注场景,对动作与表情细节建模更精准。
无词典手语翻译(Gloss-free SLT)将手语视频直接转为口语句子,避免昂贵的词典标注,但需精细建模手部、身体和面部线索。现有方法多采用单模态或弱融合特征,性能受限。本文提出ViPo-MLLM框架,整合时空RGB图像与人体姿态特征。专用编码器捕捉模态内动态,交叉注意力机制建模跨模态长程依赖。融合表征通过结构化提示输入,并由经过对比学习与语言建模训练的大型语言模型处理。在PHOENIX14T和CSL-Daily数据集上的实验表明,该模型达到新最优性能。此外,其表现媲美基于词典的手语识别方法,验证了所提姿态线索与跨模态注意力的有效性。
原文摘要 · Abstract (English)
Gloss-free Sign Language Translation (SLT) translates sign language videos into spoken-language sentences without gloss annotations, avoiding costly labeling but requiring fine-grained modeling of hands, body, and facial cues. Existing methods often use single-modality or weakly fused features, limiting performance. We propose ViPo-MLLM, a framework that integrates spatio-temporal RGB and human pose features. Dedicated encoders model intra-modal dynamics and cross-modal attention captures long-range dependencies. The fused representation is conditioned with a structured prompt and processed by an LLM trained with contrastive and language modeling objectives. The proposed model was evaluated on the PHOENIX14T and CSL-Daily datasets and achieved new state-of-the-art results on both datasets. Moreover, the ViPo-MLLM model attained competitive performance compared to gloss-based recognition approaches, confirming the effectiveness of the proposed pose cues and cross-modal attention mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。