arXiv:2604.18034cs.CLcs.CV2026-04

用多层级偏好优化提升手语骨架翻译的精准度

SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation

论文配图:SignDPO: Multi-level Direct Preference Optimisation for Skeleton-based Gloss-free Sign Language Translation
图 1 · 摘自论文原文
  • 通过分层扰动构建多粒度非偏好样本
  • 自引导机制识别关键骨骼区域提升辨别力
  • 自动语言级偏好生成避免人工标注

我们提出SignDPO,一种新型多层级直接偏好优化框架,用于增强基于骨架的手语翻译对齐。尽管现有骨架模型在最大似然估计下取得进展,但仍受限于缺乏对细粒度时空特征的判别敏感性,常导致语义漂移。SignDPO将优化目标从简单序列模仿转向空间、时间与语言维度的结构化偏好对齐。其核心设计包括:1)引入分层扰动策略,自动构建全局与局部粒度的空间和时间非偏好样本;2)提出自引导机制,利用解码器交叉注意力分数识别并扰动语义显著的骨骼区域,迫使模型区分真实手语信号与结构扭曲;3)通过微调专用扰动模型实现自动化语言级偏好生成,捕捉复杂输出级失败模式而无需人工标注。在CSL-Daily、How2Sign和OpenASL三个主流基准上的大量实验表明,SignDPO持续优于当前最先进的无词素方法,甚至媲美成熟有词素方法。结果表明,多层级偏好对齐是弥合高熵骨架轨迹与离散语言语义间鸿沟的强大范式。

原文摘要 · Abstract (English)

We present SignDPO, a novel multi-level Direct Preference Optimisation (DPO) framework designed to enhance the alignment of skeleton-based Sign Language Translation. While current skeleton-based models have made significant progress using Maximum Likelihood Estimation, they are primarily constrained by an imitation-based paradigm that lacks discriminative sensitivity to the fine-grained spatio-temporal nuances of sign language, often leading to semantic drift. To address this, SignDPO shifts the optimisation goal from simple sequence mimicry to structured preference alignment across spatial, temporal, and linguistic dimensions. Our framework involves three key designs. First, we introduce a hierarchical perturbation strategy to construct spatial and temporal non-preferred samples at both global and local granularities automatically. Second, we propose a self-guiding mechanism that leverages decoder cross-attention scores to identify and perturb semantically salient skeletal regions, forcing the model to distinguish genuine sign signals from structural distortions. Third, we establish an automated language-level preference generator by fine-tuning a dedicated perturbation model, capturing complex output-level failure modes without manual annotation. Extensive experiments on three widely adopted benchmarks, CSL-Daily, How2Sign, and OpenASL, demonstrate that SignDPO consistently outperforms state-of-the-art gloss-free methods and even rivals established gloss-based ones. Our results suggest that multi-level preference alignment is a powerful paradigm for bridging the gap between high-entropy skeletal trajectories and discrete linguistic semantics.

手语翻译偏好优化骨架模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。