跨语言手语翻译新框架,通过动作级知识迁移提升泛化能力
SIGNET: Motion-Level Knowledge Transfer for Cross-Language Sign Language Translation

- 基于注意力与手部先验的融合机制,动态选择多预训练模型专家
- 在四个数据集上达到顶尖翻译性能,跨语言迁移效果显著
- 适合研究跨语言手语识别与迁移学习的学者参考
手语翻译因高时空复杂性、长序列及多发音器官建模需求而极具挑战,且通常依赖词素标注。现有方法多针对单一数据集或语言,难以扩展,忽视了不同手语间的潜在关联。本文提出SIGNET框架,实现跨语言手语翻译中的动作级知识迁移。核心思想是:尽管手语语法与词汇存在差异,但预训练模型捕捉到的动作级视觉模式可跨数据集和语言复用。SIGNET通过基于注意力的手部先验聚合机制,整合多个预训练手语主干网络,并引导门控融合网络动态选择最相关专家。在How2Sign、Phoenix14T、CSL-Daily和MeineDGS四个基准上的实验表明,SIGNET达到当前最优翻译性能;同时在WLASL数据集上也优于以往手势识别方法。
原文摘要 · Abstract (English)
Sign language translation (SLT) remains challenging due to its high spatio-temporal complexity, long sequences, and the need to model multiple articulators without relying on gloss annotations. Existing approaches are typically tailored to individual datasets or languages and struggle to scale, while overlooking the relationships between sign languages that could inform more effective cross-lingual transfer. We present \textbf{SIGNET}, a framework that enables motion-level knowledge transfer for cross-language sign language translation. Our key insight is that, although sign languages differ in grammar and lexicon, pretrained models capture motion-level visual patterns that can be reused across datasets and languages. \textbf{SIGNET} integrates multiple pretrained sign language backbones through an attention-based, hand-prior aggregation mechanism that guides a gated fusion network in dynamically selecting the most relevant experts. Comprehensive experiments on four benchmarks (How2Sign, Phoenix14T, CSL-Daily, and MeineDGS) demonstrate state-of-the-art translation performance, and \textbf{SIGNET} also surpasses prior methods on WLASL for sign language recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。