不依赖手语标签,直接从视频翻译手语,提升准确性。
Beyond Gloss: A Hand-Centric Framework for Gloss-Free Sign Language Translation
- 用视频大模型捕捉手部动态,生成细粒度时序描述。
- 在Phoenix14T和CSL-Daily上达到当前最好效果。
- 适合做无标签手语翻译的开发者与研究者。
手语翻译(SLT)需弥合视觉与语言模态间的差距,同时捕捉手形与动作的细微差异。为此,我们提出新颖的无手语标签框架BeyondGloss,利用视频大语言模型(VideoLLMs)的时空推理能力。针对现有VideoLLMs难以精细建模长视频的问题,我们设计方法生成细粒度、时序感知的手部运动文本描述。通过对比对齐模块,在预训练中将这些描述与视频特征对齐,促使模型聚焦手部时序动态,更有效区分手势。为增强手部特异性表示,我们从HaMeR模型蒸馏细粒度特征。此外,引入符号视频表示与目标语言嵌入间的对比损失,以缩小预训练阶段的模态差距。BeyondGloss在Phoenix14T与CSL-Daily基准上取得当前最优性能。代码将于论文接受后公开。
原文摘要 · Abstract (English)
Sign Language Translation (SLT) is a challenging task that requires bridging the modality gap between visual and linguistic information while capturing subtle variations in hand shapes and movements. To address these challenges, we introduce \textbf{BeyondGloss}, a novel gloss-free SLT framework that leverages the spatio-temporal reasoning capabilities of Video Large Language Models (VideoLLMs). Since existing VideoLLMs struggle to model long videos in detail, we propose a novel approach to generate fine-grained, temporally-aware textual descriptions of hand motion. A contrastive alignment module aligns these descriptions with video features during pre-training, encouraging the model to focus on hand-centric temporal dynamics and distinguish signs more effectively. To further enrich hand-specific representations, we distill fine-grained features from HaMeR. Additionally, we apply a contrastive loss between sign video representations and target language embeddings to reduce the modality gap in pre-training. \textbf{BeyondGloss} achieves state-of-the-art performance on the Phoenix14T and CSL-Daily benchmarks, demonstrating the effectiveness of the proposed framework. We will release the code upon acceptance of the paper.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。