用分段感知的视觉编码减少视频序列长度,提升手语翻译效率
SAGE: Segment-Aware Gloss-Free Encoding for Token-Efficient Sign Language Translation
- 基于手语分段生成带语义的离散视觉标记,压缩输入序列
- 在PHOENIX14T上超越现有方法,序列长度减少50%,内存降低2.67倍
- 无需词素标注,适合大规模手语数据的高效训练与部署
无词素手语翻译(SLT)进展迅速,已无需依赖词素标注即可取得优异性能。然而,这些成果常伴随模型复杂度上升和高计算开销,制约其在大规模手语数据集上的可扩展性。本文提出一种分段感知的视觉标记化框架,利用手语分段将连续视频转化为具语义信息的离散视觉标记,使输入序列长度相比先前方法最多减少50%,从而实现最高达2.67倍的内存节省,并显著提升在更大数据集上的可扩展性。为弥合视觉与语言模态差异,引入标记到标记的对比对齐目标,以及双层监督机制,对齐语言嵌入与中间隐藏状态,实现细粒度跨模态对齐,且无需词素级监督。在PHOENIX14T基准测试中,本方法显著优于当前最优模型,同时大幅缩短序列长度。进一步实验表明,在相近序列长度下仍优于前序工作,验证了所提标记化与对齐策略的有效性。
原文摘要 · Abstract (English)
Gloss-free Sign Language Translation (SLT) has advanced rapidly, achieving strong performances without relying on gloss annotations. However, these gains have often come with increased model complexity and high computational demands, raising concerns about scalability, especially as large-scale sign language datasets become more common. We propose a segment-aware visual tokenization framework that leverages sign segmentation to convert continuous video into discrete, sign-informed visual tokens. This reduces input sequence length by up to 50% compared to prior methods, resulting in up to 2.67x lower memory usage and better scalability on larger datasets. To bridge the visual and linguistic modalities, we introduce a token-to-token contrastive alignment objective, along with a dual-level supervision that aligns both language embeddings and intermediate hidden states. This improves fine-grained cross-modal alignment without relying on gloss-level supervision. Our approach notably exceeds the performance of state-of-the-art methods on the PHOENIX14T benchmark, while significantly reducing sequence length. Further experiments also demonstrate our improved performance over prior work under comparable sequence-lengths, validating the potential of our tokenization and alignment strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。