构建首个多视角手语识别数据集,强调3D几何信息对识别的关键作用
The NGT200 Dataset: Geometric Multi-View Isolated Sign Recognition
- 提出基于3D几何对称性的手语表征方法
- 多视角识别性能提升8%-22%
- 适合手语识别与具身智能研究者
手语处理(SLP)为语言技术的包容性未来奠定基础,但其在实际应用中仍面临诸多挑战。本文聚焦多视角孤立手语识别(MV-ISR),强调3D感知与几何信息在SLP系统中的关键作用。我们提出NGT200数据集,一个新颖的时空多视角基准,将MV-ISR与单视角识别(SV-ISR)区分开来。实验表明,合成数据有效提升性能,并提出将手语表征条件化于手语内在的空间对称性。采用SE(2)等变模型,使多视角识别性能相比基线提升8%-22%。
原文摘要 · Abstract (English)
Sign Language Processing (SLP) provides a foundation for a more inclusive future in language technology; however, the field faces several significant challenges that must be addressed to achieve practical, real-world applications. This work addresses multi-view isolated sign recognition (MV-ISR), and highlights the essential role of 3D awareness and geometry in SLP systems. We introduce the NGT200 dataset, a novel spatio-temporal multi-view benchmark, establishing MV-ISR as distinct from single-view ISR (SV-ISR). We demonstrate the benefits of synthetic data and propose conditioning sign representations on spatial symmetries inherent in sign language. Leveraging an SE(2) equivariant model improves MV-ISR performance by 8%-22% over the baseline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。