arXiv:2409.15284cs.CVcs.CL2024-09被引 4

构建首个多视角手语识别数据集,强调3D几何信息对识别的关键作用

The NGT200 Dataset: Geometric Multi-View Isolated Sign Recognition

  • 提出基于3D几何对称性的手语表征方法
  • 多视角识别性能提升8%-22%
  • 适合手语识别与具身智能研究者

手语处理(SLP)为语言技术的包容性未来奠定基础,但其在实际应用中仍面临诸多挑战。本文聚焦多视角孤立手语识别(MV-ISR),强调3D感知与几何信息在SLP系统中的关键作用。我们提出NGT200数据集,一个新颖的时空多视角基准,将MV-ISR与单视角识别(SV-ISR)区分开来。实验表明,合成数据有效提升性能,并提出将手语表征条件化于手语内在的空间对称性。采用SE(2)等变模型,使多视角识别性能相比基线提升8%-22%。

原文摘要 · Abstract (English)

Sign Language Processing (SLP) provides a foundation for a more inclusive future in language technology; however, the field faces several significant challenges that must be addressed to achieve practical, real-world applications. This work addresses multi-view isolated sign recognition (MV-ISR), and highlights the essential role of 3D awareness and geometry in SLP systems. We introduce the NGT200 dataset, a novel spatio-temporal multi-view benchmark, establishing MV-ISR as distinct from single-view ISR (SV-ISR). We demonstrate the benefits of synthetic data and propose conditioning sign representations on spatial symmetries inherent in sign language. Leveraging an SE(2) equivariant model improves MV-ISR performance by 8%-22% over the baseline.

手语识别多视角3D几何数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。