用时序对齐提升手语跨语言识别,显著改善美国手语识别效果。
Cross-Sign Language Transfer Learning Using Domain Adaptation with Multi-scale Temporal Alignment
- 引入多尺度时间关系对齐的领域自适应方法,优化跨语言手语特征匹配。
- 在ASL识别上表现优于传统迁移学习,短时序特征对齐效果更佳。
- 实验证明RGB模式优于光流,适合多数手语识别任务。
手语是听力障碍者的重要沟通方式,但全球100多种手语的识别资源严重不足。为此,我们提出基于迁移学习与领域自适应方法TA3N的手语识别方案,采用时间关系网络(TRN)模块对多尺度时间关系进行对齐。实验表明,领域自适应在提升美国手语(ASL)识别性能方面优于基于神经网络的迁移学习,尤其在对齐源域与目标域的短时序特征时效果显著。除使用RGB图像外,还测试了光流模式,结果发现绝大多数情况下RGB表现更优。本研究旨在提升依赖手语为主要交流方式人群的沟通可及性。
原文摘要 · Abstract (English)
Sign language serves as a vital means of communication for individuals with hearing impairments, yet recognition resources for the over 100 distinct sign languages are severely lacking. In response, we present our work on sign language recognition using transfer learning and the domain adaptation method TA3N, which utilizes the Temporal Relational Network (TRN) module for aligning multi-scale temporal relations. Our findings highlight the superior performance of Domain Adaptation to neural network-based transfer learning, particularly in improving recognition of American Sign Language (ASL). Our research also identifies the effectiveness of aligning shorter-term temporal features between source and target domains. In addition to using RGB, we conducted experiments using Optical Flow mode for the sign language samples, ultimately determining that RGB outperforms Optical Flow in the majority of cases. Our work aims to improve accessibility and communication for individuals who rely on sign language as their primary mode of communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。