针对手语识别的签别差异和新句式泛化难题,提出双架构模型提升性能。
A Signer-Invariant Conformer and Multi-Scale Fusion Transformer for Continuous Sign Language Recognition
- 设计签别无关的Conformer,融合卷积与自注意力学习鲁棒表征。
- 多尺度融合Transformer在新句子任务上达到47.78%词错误率。
- 适用于手语识别研究者及跨领域动作识别任务。
连续手语识别(CSLR)面临显著的签别差异和对新句子结构泛化能力差的问题。传统方法难以有效应对。为此,本文提出双架构框架:针对签别无关(SI)挑战,设计签别无关的Conformer,结合卷积与多头自注意力,从基于姿态的骨骼关键点中学习鲁棒的签别无关表示;针对未见句子(US)任务,提出多尺度融合Transformer,采用新型双路径时序编码器,捕捉细粒度姿势动态,增强对新颖语法结构的理解能力。在具有挑战性的Isharah-1000数据集上的实验建立了新的基准。所提Conformer在SI挑战中实现13.07%的词错误率(WER),较最先进方法降低13.53%;在US任务中,该Transformer模型达到47.78%的WER,优于此前工作。在SignEval 2025 CSLR挑战赛中,团队在US任务中获第2名,在SI任务中获第4名,验证了模型有效性。结果表明,为特定挑战定制任务专用网络可显著提升性能,并确立了新的研究基线。代码已开源。
原文摘要 · Abstract (English)
Continuous Sign Language Recognition (CSLR) faces multiple challenges, including significant inter-signer variability and poor generalization to novel sentence structures. Traditional solutions frequently fail to handle these issues efficiently. For overcoming these constraints, we propose a dual-architecture framework. For the Signer-Independent (SI) challenge, we propose a Signer-Invariant Conformer that combines convolutions with multi-head self-attention to learn robust, signer-agnostic representations from pose-based skeletal keypoints. For the Unseen-Sentences (US) task, we designed a Multi-Scale Fusion Transformer with a novel dual-path temporal encoder that captures both fine-grained posture dynamics, enabling the model's ability to comprehend novel grammatical compositions. Experiments on the challenging Isharah-1000 dataset establish a new standard for both CSLR benchmarks. The proposed conformer architecture achieves a Word Error Rate (WER) of 13.07% on the SI challenge, a reduction of 13.53% from the state-of-the-art. On the US task, the transformer model scores a WER of 47.78%, surpassing previous work. In the SignEval 2025 CSLR challenge, our team placed 2nd in the US task and 4th in the SI task, demonstrating the performance of these models. The findings validate our key hypothesis: that developing task-specific networks designed for the particular challenges of CSLR leads to considerable performance improvements and establishes a new baseline for further research. The source code is available at: https://github.com/rezwanh001/MSLR-Pose86K-CSLR-Isharah.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。