arXiv:2507.10306cs.CV2025-07中稿 · 9th Workshop on Si…被引 1

用双视觉编码器对比预训练,实现无需手语词的自动翻译。

Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation

  • 双视觉编码器联合对齐视频与文本特征,提升表示质量。
  • 在Phoenix-2014T上达到最高BLEU-4分数,优于现有无词典方法。
  • 适合研究手语自动翻译与多模态预训练的开发者参考。

手语翻译(SLT)旨在将手语视频转换为口语或书面文本。早期系统依赖代价高昂且不完整的手语词(gloss)标注作为中间监督信号。本文提出一种两阶段、双视觉编码器框架,用于无词典手语翻译,利用对比视觉-语言预训练。预训练阶段,采用两个互补的视觉骨干网络,其输出通过对比目标与句子级文本嵌入共同对齐。下游任务中,融合视觉特征输入编码器-解码器模型。在Phoenix-2014T基准测试中,该双编码器架构持续优于单流变体,成为现有无词典方法中BLEU-4得分最高的方案。

原文摘要 · Abstract (English)

Sign Language Translation (SLT) aims to convert sign language videos into spoken or written text. While early systems relied on gloss annotations as an intermediate supervision, such annotations are costly to obtain and often fail to capture the full complexity of continuous signing. In this work, we propose a two-phase, dual visual encoder framework for gloss-free SLT, leveraging contrastive visual-language pretraining. During pretraining, our approach employs two complementary visual backbones whose outputs are jointly aligned with each other and with sentence-level text embeddings via a contrastive objective. During the downstream SLT task, we fuse the visual features and input them into an encoder-decoder model. On the Phoenix-2014T benchmark, our dual encoder architecture consistently outperforms its single stream variants and achieves the highest BLEU-4 score among existing gloss-free SLT approaches.

手语翻译对比学习双编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。