arXiv:2508.01791cs.CVcs.AI2025-08被引 1

用数据驱动方法提升阿拉伯手语识别准确率,3秒读懂核心思路。

CSLRConformer: A Data-Centric Conformer Approach for Continuous Arabic Sign Language Recognition on the Isharah Datase

  • 基于探索性数据分析筛选关键手势点,构建高效特征
  • 预处理融合DBSCAN去噪与空间归一化,提升数据质量
  • 改造语音模型Conformer用于手语识别,实现新基准

连续手语识别(CSLR)面临流畅的跨手势过渡、无时间边界和共发音效应等挑战。本文针对签名者无关识别问题,提出一种以数据为中心的方法,包含系统性特征工程、稳健预处理流程和优化模型架构。核心贡献包括:基于探索性数据分析(EDA)的特征选择,以提取有通信意义的关键点;结合DBSCAN的异常值过滤与空间归一化的严格预处理流程;以及新型的CSLRConformer架构。该架构改编自面向语音识别的混合CNN-Transformer结构,能有效建模局部时序依赖与全局序列上下文,适用于手语的时空动态特性。在开发集上取得5.60%的词错误率(WER),测试集为12.01%,在ICCV 2025 MSLR 2025挑战赛中排名第三。研究验证了跨领域架构迁移的有效性,表明原用于语音识别的Conformer可成功应用于基于关键点的手语识别,建立新基准。

原文摘要 · Abstract (English)

The field of Continuous Sign Language Recognition (CSLR) poses substantial technical challenges, including fluid inter-sign transitions, the absence of temporal boundaries, and co-articulation effects. This paper, developed for the MSLR 2025 Workshop Challenge at ICCV 2025, addresses the critical challenge of signer-independent recognition to advance the generalization capabilities of CSLR systems across diverse signers. A data-centric methodology is proposed, centered on systematic feature engineering, a robust preprocessing pipeline, and an optimized model architecture. Key contributions include a principled feature selection process guided by Exploratory Data Analysis (EDA) to isolate communicative keypoints, a rigorous preprocessing pipeline incorporating DBSCAN-based outlier filtering and spatial normalization, and the novel CSLRConformer architecture. This architecture adapts the hybrid CNN-Transformer design of the Conformer model, leveraging its capacity to model local temporal dependencies and global sequence context; a characteristic uniquely suited for the spatio-temporal dynamics of sign language. The proposed methodology achieved a competitive performance, with a Word Error Rate (WER) of 5.60% on the development set and 12.01% on the test set, a result that secured a 3rd place ranking on the official competition platform. This research validates the efficacy of cross-domain architectural adaptation, demonstrating that the Conformer model, originally conceived for speech recognition, can be successfully repurposed to establish a new state-of-the-art performance in keypoint-based CSLR.

手语识别Conformer数据驱动关键点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。