arXiv:2607.27614cs.CLcs.AI2026-07

DualAnchor提升手语翻译流畅性与词汇准确性,解决大模型语言先验退化问题。

DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation

论文配图:DualAnchor: Preserving Language Priors and Improving Lexical Fidelity in Gloss-Free Sign Language Translation
图 1 · 摘自论文原文
  • 引入双锚机制:词元级语言先验锚定(TPA)保持语言连贯性
  • 通过熵正则部分最优传输(OTA)实现视觉与文本软对齐,减少词汇错误
  • 在PHOENIX-2014T和CSL-Daily上表现优越,适合追求高精度手语翻译的研究者

近期大型语言模型(LLMs)的发展推动了手语翻译(SLT)任务采用LLM作为文本骨干。然而,现有基于LLM的SLT方法常削弱而非利用其语言先验,导致翻译不连贯,我们称之为语言先验退化。同时,现有方法通常在句子层面进行视频与文本对齐,无法保证细粒度词汇准确,造成词汇保真度差距。为此,我们提出DualAnchor——一种无词签的基于LLM的手语翻译训练框架,通过两个互补锚点实现语言流畅与视觉忠实生成。词元级先验锚定(TPA)通过在每一步解码时将多模态解码器正则化至冻结LLM在相同自回归前缀下的下一个词分布,以保留语言先验。最优传输对齐(OTA)将视觉-文本匹配建模为熵正则化部分最优传输,通过Sinkhorn优化在余弦代价下诱导视觉词元与文本内容词元之间的软对齐。DualAnchor在PHOENIX-2014T和CSL-Daily数据集上均取得优异性能。针对性分析表明,两项机制协同增效:TPA提升流畅性,而OTA降低细粒度词汇错误。

原文摘要 · Abstract (English)

Recent advances in large language models (LLMs) have led sign language translation (SLT), the task of converting sign-language videos into spoken-language text, to increasingly adopt LLMs as textual backbones. However, despite their strong language modeling capabilities, existing LLM-based SLT methods often undermine rather than exploit this language prior, producing disfluent translations, a failure we term language-prior degradation. Meanwhile, existing methods typically align videos and text at the sentence level, which does not ensure accurate lexical details and creates a lexical fidelity gap. To address both issues, we propose DualAnchor, a gloss-free LLM-based SLT training framework that couples two complementary anchors for linguistically fluent and visually faithful generation. Token-level Prior Anchoring (TPA) preserves the LLM's language prior by regularizing the multimodal decoder at each decoding step toward the next-token distribution of a frozen LLM conditioned on the same autoregressive prefix. Optimal Transport Alignment (OTA) improves lexical fidelity by formulating visual-textual matching as entropy-regularized partial optimal transport, with Sinkhorn optimization inducing a soft alignment between visual tokens and textual content tokens under a cosine cost. DualAnchor achieves strong overall performance on both PHOENIX-2014T and CSL-Daily. Targeted analyses attribute these gains to the complementary effects of the two anchors: TPA improves fluency, whereas OTA reduces fine-grained lexical errors.

手语翻译大模型对齐机制语言先验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。