arXiv:2504.16315cs.CVcs.CL2025-04被引 5

SignX用紧凑的姿势空间实现高效连续手语识别,速度提升近50倍。

SignX: Continuous Sign Recognition in Compact Pose-Rich Latent Space

  • 将多种姿势格式融合到紧凑的潜在空间中统一表示。
  • 在连续手语识别任务上达到最新性能,比像素空间基线快近50倍。
  • 适合需要实时手语识别的应用,如无障碍通信系统。

手语数据处理复杂度高,现有方法通常通过姿态信息将手语视频翻译为词级标识符(Glosses)进行识别。本文提出SignX,一种在紧凑姿势丰富潜在空间中的连续手语识别新框架。首先,构建统一的潜在表示,将SMPLer-X、DWPose、MediaPipe、PrimeDepth和Sapiens Segmentation等多种异构姿态格式编码至紧凑且信息密集的空间。其次,训练基于ViT的视频到姿态模块,直接从原始视频中提取该潜在表示。最后,在此潜在空间中设计时序建模与序列优化方法,实现端到端手语识别,显著降低计算开销。实验表明,SignX在连续手语识别与翻译任务上达到当前最优性能,相比像素空间基线实现近50倍加速。

原文摘要 · Abstract (English)

The complexity of Sign Language (SL) data processing brings many challenges. The current approach to recognition of SL signs aims to translate RGB sign language videos through pose information into Word-based ID Glosses, which serve to uniquely identify signs. This paper proposes SignX, a novel framework for continuous sign language recognition (SLR) in compact pose-rich latent space. First, we construct a unified latent representation that encodes heterogeneous pose formats (SMPLer-X, DWPose, Mediapipe, PrimeDepth, and Sapiens Segmentation) into a compact, information-dense space. Second, we train a ViT-based Video-to-Pose module to extract this latent representation directly from raw videos. Finally, we develop a temporal modeling and sequence refinement method that operates entirely in this latent space. This multi-stage design achieves end-to-end SLR while significantly reducing computational consumption. Experimental results demonstrate that SignX achieves SOTA accuracy on continuous SLR and Translation task, delivering nearly a 50-fold acceleration over pixel-space baselines.

手语识别姿态表示视觉模型高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。