用音位结构分解提升手语识别,小样本下效果显著
State Space Models are Effective Sign Language Learners: Exploiting Phonological Compositionality for Vocabulary-Scale Recognition
- 将手语拆解为手形、位置、动作等音位参数,构建分层表示
- 在5565个手语词的大型数据集上达到72.1%准确率,小样本提升225%
- 适合需要少样本学习或跨数据集迁移的手语研究与应用
手语识别面临灾难性扩展失败:小词汇量模型在真实规模下性能崩溃。现有方法将手势视为原子视觉模式,学习扁平表示,无法利用手语的组合结构——即由离散音位参数(手形、位置、动作、朝向)系统性重复构成。我们提出PHONSSM,通过解剖学基础的图注意力、显式正交子空间分解和原型分类,强制实现音位分解。仅使用骨架数据,在迄今最大的美式手语数据集(5,565个符号)上,取得WLASL2000上72.1%的准确率,比骨架方法领先18.4个百分点,超越多数无需视频输入的RGB方法。小样本情形下提升达225%,零样本迁移至ASL Citizen也优于监督RGB基线。词汇扩展瓶颈本质是表征学习问题,可通过模仿语言结构的组合归纳偏置解决。
原文摘要 · Abstract (English)
Sign language recognition suffers from catastrophic scaling failure: models achieving high accuracy on small vocabularies collapse at realistic sizes. Existing architectures treat signs as atomic visual patterns, learning flat representations that cannot exploit the compositional structure of sign languages-systematically organized from discrete phonological parameters (handshape, location, movement, orientation) reused across the vocabulary. We introduce PHONSSM, enforcing phonological decomposition through anatomically-grounded graph attention, explicit factorization into orthogonal subspaces, and prototypical classification enabling few-shot transfer. Using skeleton data alone on the largest ASL dataset ever assembled (5,565 signs), PHONSSM achieves 72.1% on WLASL2000 (+18.4pp over skeleton SOTA), surpassing most RGB methods without video input. Gains are most dramatic in the few-shot regime (+225% relative), and the model transfers zero-shot to ASL Citizen, exceeding supervised RGB baselines. The vocabulary scaling bottleneck is fundamentally a representation learning problem, solvable through compositional inductive biases mirroring linguistic structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。