arXiv:2412.11553cs.CV2024-12被引 4

针对手语识别数据质量差、速度不一,提出一套高效训练方案

Training Strategies for Isolated Sign Language Recognition

  • 结合图像视频增强与新损失函数,提升模型对动作时序的感知
  • 在WLASL和Slovo数据集上达到当前最优性能
  • 方案通用性强,适配多种模型架构与数据集

准确识别和解读手语对提升聋人及重听人士的沟通无障碍性至关重要。然而,当前孤立手语识别(ISLR)方法常面临数据质量低、手势速度差异大等挑战。本文提出一套完整的ISLR模型训练流程,旨在适应手语领域的独特特征与限制。该流程采用精心选择的图像与视频增强技术,以应对低质量数据和不同手势速度问题。引入额外回归头并结合IoU平衡分类损失,增强模型对动作起止时间的感知,简化时序信息捕捉。大量实验表明,所提训练流程可轻松适配不同数据集与网络结构。消融研究显示,每一组件均有助于更好地考虑ISLR任务特性。所提策略在多个ISLR基准上显著提升识别性能,并在WLASL与Slovo数据集上取得当前最佳结果。

原文摘要 · Abstract (English)

Accurate recognition and interpretation of sign language are crucial for enhancing communication accessibility for deaf and hard of hearing individuals. However, current approaches of Isolated Sign Language Recognition (ISLR) often face challenges such as low data quality and variability in gesturing speed. This paper introduces a comprehensive model training pipeline for ISLR designed to accommodate the distinctive characteristics and constraints of the Sign Language (SL) domain. The constructed pipeline incorporates carefully selected image and video augmentations to tackle the challenges of low data quality and varying sign speeds. Including an additional regression head combined with IoU-balanced classification loss enhances the model's awareness of the gesture and simplifies capturing temporal information. Extensive experiments demonstrate that the developed training pipeline easily adapts to different datasets and architectures. Additionally, the ablation study shows that each proposed component expands the potential to consider ISLR task specifics. The presented strategies enhance recognition performance across various ISLR benchmarks and achieve state-of-the-art results on the WLASL and Slovo datasets.

手语识别数据增强时序建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。