arXiv:2604.22374cs.CL2026-04ACL被引 1

提出新方法提升无词签手语翻译的跨模态对齐效果

Selective Contrastive Learning For Gloss Free Sign Language Translation

论文配图:Selective Contrastive Learning For Gloss Free Sign Language Translation
图 1 · 摘自论文原文
  • 基于相似性动态筛选负样本,避免无效对比
  • 渐进式构建批次,强化难例负样本监督
  • 适合研究手语翻译与跨模态学习的学者

手语翻译(SLT)将连续手语视频转换为口语文本,但因视觉手势与书面语言之间存在固有模态差异,尤其在无词签(gloss-free)场景下仍具挑战。现有SLT系统越来越多地采用类似CLIP的视觉-语言预训练(VLP)进行跨模态对齐,但随机的批内对比仅提供少量依赖批次的负样本,且常将语义相近甚至相同的对误标为负样本,引入噪声和不一致的对齐监督。本文首次开展基于轨迹的分析,追踪训练过程中负样本视频-文本相似度变化。结果表明,仅有小部分负样本表现出持续远离的期望行为,其余负样本呈现异质性且相似度往往不下降,说明随机负样本常不具备有效对齐信息。受此启发,我们提出选择性对比学习框架SCL-SLT,其核心为配对选择(PS)策略:利用参考检查点的相似性动态评分候选负样本,并通过渐进式课程学习构建小批量,逐步强化困难负样本的对比监督,从而减少噪声或语义无效负样本的影响。

原文摘要 · Abstract (English)

Sign language translation (SLT) converts continuous sign videos into spoken-language text, yet it remains challenging due to the intrinsic modality mismatch between visual signs and written text, particularly in gloss-free settings. Recent SLT systems increasingly adopt CLIP-like Vision-Language pretraining (VLP) for cross-modal alignment, but the random in-batch contrast provides few, batch-dependent negatives and may mislabel semantically similar (or even identical) pairs as negatives, introducing noisy and potentially inconsistent alignment supervision. In this work, we first conduct a preliminary trajectory-based analysis that tracks negative video-text similarity over training. The results show that only a small subset of negatives exhibits the desired behavior of being consistently pushed away, while the remaining negatives display heterogeneous and often non-decreasing similarity dynamics, suggesting that random in-batch negatives are frequently uninformative for effective alignment. Inspired by this, we propose Selective Contrastive Learning for SLT (SCL-SLT) with a Pair Selection (PS) strategy. PS scores candidate negatives using similarity dynamics from reference checkpoints and constructs mini-batches via a curriculum that progressively emphasizes more challenging negatives, thereby strengthening contrastive supervision while reducing the influence of noisy or semantically invalid negatives.

手语翻译对比学习跨模态对齐视觉语言预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。