arXiv:2512.07273cs.CV2025-12被引 2

用强化学习提升手语翻译,让模型更懂动作细节和句子完整。

RVLF: A Reinforcing Vision-Language Framework for Gloss-Free Sign Language Translation

  • 融合骨骼运动与视觉特征,构建手语专用视觉语言模型
  • 引入基于GRPO的优化策略,提升翻译准确性和句子完整性
  • 无需外部数据预训练,多数据集显著提升翻译质量

无词义标注的手语翻译受限于两个关键问题:难以捕捉细微视觉线索的表征不足,以及现有大模型方法在句级语义对齐上的偏差。为此,我们提出三阶段强化视觉-语言框架(RVLF)。首先构建专用于手语的大型视觉语言模型(LVLM),融合骨骼运动信息与DINOv2提取的语义丰富视觉特征,并通过指令微调获得强基线模型SLT-SFT;其次引入基于GRPO的优化策略,利用包含翻译保真度(BLEU)与句子完整性(ROUGE)的奖励函数微调模型,得到优化模型SLT-GRPO。该框架无需任何外部大规模手语数据预训练,在无词义标注设定下实现显著提升:在CSL-Daily、PHOENIX-2014T、How2Sign和OpenASL数据集上,BLEU-4分数分别提升+5.1、+1.11、+1.4和+1.61。据我们所知,这是首个将GRPO引入手语翻译的工作。大量实验与消融分析验证了基于GRPO的优化在提升翻译质量与语义一致性方面的有效性。

原文摘要 · Abstract (English)

Gloss-free sign language translation (SLT) is hindered by two key challenges: **inadequate sign representation** that fails to capture nuanced visual cues, and **sentence-level semantic misalignment** in current LLM-based methods, which limits translation quality. To address these issues, we propose a three-stage **r**einforcing **v**ision-**l**anguage **f**ramework (**RVLF**). We build a large vision-language model (LVLM) specifically designed for sign language, and then combine it with reinforcement learning (RL) to adaptively enhance translation performance. First, for a sufficient representation of sign language, RVLF introduces an effective semantic representation learning mechanism that fuses skeleton-based motion cues with semantically rich visual features extracted via DINOv2, followed by instruction tuning to obtain a strong SLT-SFT baseline. Then, to improve sentence-level semantic misalignment, we introduce a GRPO-based optimization strategy that fine-tunes the SLT-SFT model with a reward function combining translation fidelity (BLEU) and sentence completeness (ROUGE), yielding the optimized model termed SLT-GRPO. Our conceptually simple framework yields substantial gains under the gloss-free SLT setting without pre-training on any external large-scale sign language datasets, improving BLEU-4 scores by +5.1, +1.11, +1.4, and +1.61 on the CSL-Daily, PHOENIX-2014T, How2Sign, and OpenASL datasets, respectively. To the best of our knowledge, this is the first work to incorporate GRPO into SLT. Extensive experiments and ablation studies validate the effectiveness of GRPO-based optimization in enhancing both translation quality and semantic consistency.

手语翻译视觉语言模型强化学习多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。