用循环注意力与几何对齐,提升手语识别精度
LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition
- 通过循环重构隐状态,逐步优化动作理解
- 在两个数据集上达到新最好效果,层参数更少
- 适合关注动作细粒度建模的视觉算法研究者
基于骨架的孤立手语识别(ISLR)需要对多尺度运动细节有精细理解,从细微手指动作到整体身体动态。现有方法多依赖深层前馈网络,虽提升模型容量,但缺乏递归优化机制与结构化表征能力。本文提出LA-Sign,一种具有几何感知对齐的循环变换器框架。不通过堆叠更深层,而是利用重复访问隐状态实现深度,共享参数下逐次精炼运动理解。为进一步正则化该过程,提出几何感知对比目标,将骨架与文本特征映射至自适应双曲空间,促进多尺度语义组织。研究三种循环设计及多种几何流形,表明编码器-解码器循环结合自适应Poincaré对齐性能最优。在WLASL和MSASL基准上的大量实验显示,LA-Sign在使用更少独立层的情况下达成当前最佳性能,验证了递归隐状态精炼与几何感知表示学习在手语识别中的有效性。
原文摘要 · Abstract (English)
Skeleton-based isolated sign language recognition (ISLR) demands fine-grained understanding of articulated motion across multiple spatial scales, from subtle finger movements to global body dynamics. Existing approaches typically rely on deep feed-forward architectures, which increase model capacity but lack mechanisms for recurrent refinement and structured representation. We propose LA-Sign, a looped transformer framework with geometry-aware alignment for ISLR. Instead of stacking deeper layers, LA-Sign derives its depth from recurrence, repeatedly revisiting latent representations to progressively refine motion understanding under shared parameters. To further regularise this refinement process, we present a geometry-aware contrastive objective that projects skeletal and textual features into an adaptive hyperbolic space, encouraging multi-scale semantic organisation. We study three looping designs and multiple geometric manifolds, demonstrating that encoder-decoder looping combined with adaptive Poincare alignment yields the strongest performance. Extensive experiments on WLASL and MSASL benchmarks show that LA-Sign achieves state-of-the-art results while using fewer unique layers, highlighting the effectiveness of recurrent latent refinement and geometry-aware representation learning for sign language recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。