arXiv:2601.03549cs.CVcs.CL2026-01中稿 · EMNLP

让手语翻译更懂表情,提升语义准确性

FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation

  • 用面部动态增强手语手势特征,双向融合互补信息
  • 在PHOENIX14T和CSL-Daily上达到当前最优无词典性能
  • 特别擅长处理依赖表情区分含义的手语表达

手语翻译是一项需联合建模手势与非手势信号的跨模态任务。现有无词典方法虽能捕捉手势动态,但常忽视面部表情——其在语法和消歧中起关键作用。当不同概念具有相似手势时,忽略表情会导致语义退化。为此,我们提出FEA-SLT(面部表情感知手语翻译)框架,通过领域迁移的面部编码器提取表情敏感表征,并利用语言学驱动的面部-手势融合模块(FEAF),以双向调制机制捕捉两者间的相互依赖,提升句法准确率。在PHOENIX14T和CSL-Daily数据集上的实验表明,该方法在无词典条件下取得当前最优BLEU得分;针对性分析也验证了其对表情敏感语句的翻译改进。代码已开源。

原文摘要 · Abstract (English)

Sign Language Translation (SLT) is a challenging cross-modal task requiring joint modeling of manual articulations and non-manual signals. Existing gloss-free SLT methods effectively capture gestural dynamics but often underutilize facial expressions, which play crucial grammatical and disambiguating roles. This limitation can cause semantic degradation when distinct concepts share similar manual configurations. To address this issue, we propose FEA-SLT (**F**acial-**E**xpression-**A**ware **S**ign **L**anguage **T**ranslation), a gloss-free end-to-end framework that uses facial dynamics to provide complementary cues to manual signals. FEA-SLT employs a domain-transferred facial encoder to extract expression-sensitive representations and integrates them with manual features through a linguistically motivated *Facial-Expression-Aware Fusion* (FEAF) module. FEAF captures reciprocal dependencies between manual and facial channels via bidirectional modulation, enhancing syntactic fidelity. Experiments on PHOENIX14T and CSL-Daily show that FEA-SLT achieves state-of-the-art BLEU performance among gloss-free methods, while targeted analyses support improved translation of facial-sensitive utterances. Code is available at https://github.com/TuGuobin/FEA-SLT.

手语翻译面部表情端到端多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。