用双视角图卷积提升手语识别准确率
Skeleton-based sign language recognition using a dual-stream spatio-temporal dynamic graph convolutional network
- 分两路分别建模手势形状与运动轨迹
- 在三个数据集上达到93.7%~99.8%准确率
- 适合做轻量级手语识别系统开发
孤立手语识别面临形态相似但语义不同的手势挑战,根源在于手形与运动轨迹的复杂耦合。现有方法多依赖单一参考帧,难以解决几何模糊性。本文提出双参考、双流架构Dual-SignLanguageNet(DSLNet),将手势形态与轨迹解耦并在互补坐标系中建模:以腕部为中心的拓扑感知图卷积捕捉视图不变的手形特征;以面部为中心的弗林格几何编码器捕获上下文感知的运动轨迹。两者通过几何驱动的最优传输机制融合。DSLNet在挑战性数据集WLASL-100、WLASL-300和LSA64上分别取得93.70%、89.97%和99.79%的准确率,参数量显著少于对比模型。
原文摘要 · Abstract (English)
Isolated Sign Language Recognition (ISLR) is challenged by gestures that are morphologically similar yet semantically distinct, a problem rooted in the complex interplay between hand shape and motion trajectory. Existing methods, often relying on a single reference frame, struggle to resolve this geometric ambiguity. This paper introduces Dual-SignLanguageNet (DSLNet), a dual-reference, dual-stream architecture that decouples and models gesture morphology and trajectory in separate, complementary coordinate systems. The architecture processes these streams through specialized networks: a topology-aware graph convolution models the view-invariant shape from a wrist-centric frame, while a Finsler geometry-based encoder captures the context-aware trajectory from a facial-centric frame. These features are then integrated via a geometry-driven optimal transport fusion mechanism. DSLNet sets a new state-of-the-art, achieving 93.70%, 89.97%, and 99.79% accuracy on the challenging WLASL-100, WLASL-300, and LSA64 datasets, respectively, with significantly fewer parameters than competing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。