统一手势语言理解框架,提升跨任务性能
Uni-Sign: Toward Unified Sign Language Understanding at Scale
- 将下游任务统一为手势翻译任务,消除预训练与微调差距
- 基于1985小时中文手语数据集,实现大规模生成式预训练
- 融合姿态与视觉信息,提升准确率并降低计算开销
手势语言预训练因能提升多种手势语言理解(SLU)任务表现而受到关注。然而现有方法在预训练与微调之间存在鸿沟,导致效果不佳。为此,我们提出Uni-Sign,一种通过大规模生成式预训练和新型微调范式消除该鸿沟的统一预训练框架。首先,我们构建了包含1,985小时视频与文本标注的大型中文手语数据集CSL-News,支持高效的大规模预训练。其次,Uni-Sign通过将下游任务统一为单一的手势语言翻译(SLT)任务,在微调阶段确保预训练知识的无缝迁移。此外,引入先验引导融合(PGF)模块与评分感知采样策略,有效融合姿态与RGB信息,缓解关键点误差并提升计算效率。在多个SLU基准上的大量实验表明,Uni-Sign在多项下游任务中达到最先进性能。数据集与代码已开源。
原文摘要 · Abstract (English)
Sign language pre-training has gained increasing attention for its ability to enhance performance across various sign language understanding (SLU) tasks. However, existing methods often suffer from a gap between pre-training and fine-tuning, leading to suboptimal results. To address this, we propose Uni-Sign, a unified pre-training framework that eliminates the gap between pre-training and downstream SLU tasks through a large-scale generative pre-training strategy and a novel fine-tuning paradigm. First, we introduce CSL-News, a large-scale Chinese Sign Language (CSL) dataset containing 1,985 hours of video paired with textual annotations, which enables effective large-scale pre-training. Second, Uni-Sign unifies SLU tasks by treating downstream tasks as a single sign language translation (SLT) task during fine-tuning, ensuring seamless knowledge transfer between pre-training and fine-tuning. Furthermore, we incorporate a prior-guided fusion (PGF) module and a score-aware sampling strategy to efficiently fuse pose and RGB information, addressing keypoint inaccuracies and improving computational efficiency. Extensive experiments across multiple SLU benchmarks demonstrate that Uni-Sign achieves state-of-the-art performance across multiple downstream SLU tasks. Dataset and code are available at github.com/ZechengLi19/Uni-Sign.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。