用视频变压器提升孟加拉手语识别准确率,小规模数据达95.5%。
Fine-Tuning Video Transformers for Word-Level Bangla Sign Language: A Comparative Analysis for Classification Tasks
- 微调VideoMAE等视频变换模型,处理孟加拉手语视频数据
- 在60类小数据集上达95.5%准确率,401类大数据集达81.04%
- 适合手语识别研究者、无障碍技术开发者参考
手语识别(SLR)旨在从图像或视频中自动识别和分类手势,将其转化为文本或语音,以提升听障群体的沟通可及性。在孟加拉国,孟加拉手语(BdSL)是许多听障人士的主要交流方式。本研究对前沿视频变换架构——VideoMAE、ViViT和TimeSformer进行微调,使用包含60个常见手势的小规模数据集BdSLW60(arXiv:2402.08635),将视频统一为30 FPS,共获得9,307条用户试录片段。为评估可扩展性与鲁棒性,模型亦在大规模数据集BdSLW401(arXiv:2503.02360,含401个手势类别)上训练。同时,对比公开数据集LSA64与WLASL的表现。采用随机裁剪、水平翻转及短边缩放等数据增强策略提升模型泛化能力。为确保各轮次评估均衡,训练集采用10折分层交叉验证,测试阶段采用未见用户U4与U8的独立数据进行签名人无关评估。结果表明,视频变换模型显著优于传统机器学习与深度学习方法。性能受数据规模、视频质量、帧分布、帧率及模型结构影响。其中,VideoMAE变体(MCG-NJU/videomae-base-finetuned-kinetics)在帧率校正后的BdSLW60上达到95.5%准确率,在BdSLW401前向手势子集上达81.04%,展现出良好的可扩展与高精度潜力。
原文摘要 · Abstract (English)
Sign Language Recognition (SLR) involves the automatic identification and classification of sign gestures from images or video, converting them into text or speech to improve accessibility for the hearing-impaired community. In Bangladesh, Bangla Sign Language (BdSL) serves as the primary mode of communication for many individuals with hearing impairments. This study fine-tunes state-of-the-art video transformer architectures -- VideoMAE, ViViT, and TimeSformer -- on BdSLW60 (arXiv:2402.08635), a small-scale BdSL dataset with 60 frequent signs. We standardized the videos to 30 FPS, resulting in 9,307 user trial clips. To evaluate scalability and robustness, the models were also fine-tuned on BdSLW401 (arXiv:2503.02360), a large-scale dataset with 401 sign classes. Additionally, we benchmark performance against public datasets, including LSA64 and WLASL. Data augmentation techniques such as random cropping, horizontal flipping, and short-side scaling were applied to improve model robustness. To ensure balanced evaluation across folds during model selection, we employed 10-fold stratified cross-validation on the training set, while signer-independent evaluation was carried out using held-out test data from unseen users U4 and U8. Results show that video transformer models significantly outperform traditional machine learning and deep learning approaches. Performance is influenced by factors such as dataset size, video quality, frame distribution, frame rate, and model architecture. Among the models, the VideoMAE variant (MCG-NJU/videomae-base-finetuned-kinetics) achieved the highest accuracies of 95.5% on the frame rate corrected BdSLW60 dataset and 81.04% on the front-facing signs of BdSLW401 -- demonstrating strong potential for scalable and accurate BdSL recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。