用手指关节角度特征提升少样本跨语言手语识别准确率
Geometry-Aware Metric Learning for Cross-Lingual Few-Shot Sign Language Recognition on Static Hand Keypoints
- 用20维关节角度替代原始坐标,天然抵抗视角、缩放等变化
- 在4种手语字母表上,少样本跨语言迁移最高提升25个百分点
- 轻量模型仅需10万参数,适合低资源手语场景
手语识别系统通常需要每种语言大量标注数据,但全球300多种手语中多数缺乏足够标注。跨语言少样本迁移——在数据丰富的源语言上预训练,仅用少量目标语言样本微调——提供了一种可扩展的替代方案。然而,传统基于坐标的关节点表示易受相机视角、手部尺度和录制条件差异引起的域偏移影响,这种偏移在少样本情况下尤为严重,因为仅用K个样本估计的类别原型对非本质变化极为敏感。本文提出一种几何感知度量学习框架,核心是基于MediaPipe静态手部关节点的紧凑20维关节约束描述子。该描述子对SO(3)旋转、平移和各向同性缩放保持不变,消除了主要的跨数据集偏移来源,显著提升了类别原型的紧密性与稳定性。在涵盖语法类型多样的四种拼写字母表(美式手语ASL、巴西手语LIBRAS、阿拉伯手语、泰式手语)上评估,所提角度特征在同域任务上相比归一化坐标基线最高提升25个百分点,并实现了冻结模型的跨语言迁移,其性能常超过同域准确率,仅使用约10^5参数的轻量MLP编码器即可达成。结果表明,不变的手部几何描述子为低资源环境下的跨语言少样本手语识别提供了可移植且高效的基础。
原文摘要 · Abstract (English)
Sign language recognition (SLR) systems typically require large labeled corpora for each language, yet the majority of the world's 300+ sign languages lack sufficient annotated data. Cross-lingual few-shot transfer, pretraining on a data-rich source language and adapting with only a handful of target-language examples, offers a scalable alternative, but conventional coordinate-based keypoint representations are susceptible to domain shift arising from differences in camera viewpoint, hand scale, and recording conditions. This shift is particularly detrimental in the few-shot regime, where class prototypes estimated from only K examples are highly sensitive to extrinsic variance. We propose a geometry-aware metric-learning framework centered on a compact 20-dimensional inter-joint angle descriptor derived from MediaPipe static hand keypoints. These angles are invariant to SO(3) rotation, translation, and isotropic scaling, eliminating the dominant sources of cross-dataset shift and yielding tighter, more stable class prototypes. Evaluated on four fingerspelling alphabets spanning typologically diverse sign languages, ASL, LIBRAS, Arabic Sign Language, and Thai Sign Language, the proposed angle features improve over normalized-coordinate baselines by up to 25 percentage points within-domain and enable frozen cross-lingual transfer that frequently exceeds within-domain accuracy, using a lightweight MLP encoder with about 10^5 parameters. These findings demonstrate that invariant hand-geometry descriptors provide a portable and effective foundation for cross-lingual few-shot SLR in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。