用姿态与分割提升手语识别效率与鲁棒性
Isolated Sign Language Recognition with Segmentation and Pose Estimation
- 融合姿态估计与图像分割提取关键视觉特征
- 降低计算开销,保持对不同使用者的识别稳定性
- 适合资源受限场景下的手语识别系统开发
大型语言模型已实现口语和书面语的自动翻译,但对依赖复杂视觉线索的美国手语(ASL)用户仍不友好。孤立手语识别(ISLR)旨在分类单个手语视频,但受限于每类手势数据稀缺、签名者差异大及计算成本高。本文提出一种新模型,通过整合(i)姿态估计管道提取手部与面部关节点坐标,(ii)分割模块分离相关视觉信息,(iii)基于ResNet-Transformer的主干网络联合建模空间与时间依赖关系,在降低计算开销的同时,保持对签名者差异的强鲁棒性。
原文摘要 · Abstract (English)
The recent surge in large language models has automated translations of spoken and written languages. However, these advances remain largely inaccessible to American Sign Language (ASL) users, whose language relies on complex visual cues. Isolated sign language recognition (ISLR) - the task of classifying videos of individual signs - can help bridge this gap but is currently limited by scarce per-sign data, high signer variability, and substantial computational costs. We propose a model for ISLR that reduces computational requirements while maintaining robustness to signer variation. Our approach integrates (i) a pose estimation pipeline to extract hand and face joint coordinates, (ii) a segmentation module that isolates relevant information, and (iii) a ResNet-Transformer backbone to jointly model spatial and temporal dependencies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。