解决手语识别中长期动作与多方向运动的难题,提升识别准确率。
OLMD: Orientation-aware Long-term Motion Decoupling for Continuous Sign Language Recognition
- 通过长时运动聚合与方向解耦,分离复杂手语动作
- 在三个数据集上实现当前最优效果,PHOENIX14 WER降低1.6%
- 适合需要高精度连续手语识别的无障碍应用
连续手语识别(CSLR)的主要挑战源于多方向性和长时运动的存在。然而,现有研究忽视了这些关键因素,严重影响了识别精度。为此,我们提出一种新框架:面向方向感知的长时运动解耦(OLMD),该框架能高效聚合长时运动,并将多方向信号解耦为易于理解的分量。具体而言,创新的长时运动聚合(LMA)模块可过滤静态冗余,自适应捕获长时运动的丰富特征;通过将复杂动作分解为水平和垂直分量,增强方向感知能力,实现双方向运动净化。此外,引入阶段内与跨阶段耦合机制,共同丰富多尺度特征,提升模型泛化能力。实验表明,OLMD在三个大规模数据集——PHOENIX14、PHOENIX14-T和CSL-Daily上均达到当前最优性能,尤其在PHOENIX14上将词错误率(WER)绝对降低1.6%。
原文摘要 · Abstract (English)
The primary challenge in continuous sign language recognition (CSLR) mainly stems from the presence of multi-orientational and long-term motions. However, current research overlooks these crucial aspects, significantly impacting accuracy. To tackle these issues, we propose a novel CSLR framework: Orientation-aware Long-term Motion Decoupling (OLMD), which efficiently aggregates long-term motions and decouples multi-orientational signals into easily interpretable components. Specifically, our innovative Long-term Motion Aggregation (LMA) module filters out static redundancy while adaptively capturing abundant features of long-term motions. We further enhance orientation awareness by decoupling complex movements into horizontal and vertical components, allowing for motion purification in both orientations. Additionally, two coupling mechanisms are proposed: stage and cross-stage coupling, which together enrich multi-scale features and improve the generalization capabilities of the model. Experimentally, OLMD shows SOTA performance on three large-scale datasets: PHOENIX14, PHOENIX14-T, and CSL-Daily. Notably, we improved the word error rate (WER) on PHOENIX14 by an absolute 1.6% compared to the previous SOTA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。