arXiv:2509.04745cs.CLcs.CV2025-09被引 1

通过语音学先验提升手语表征,增强对未见手势的泛化能力。

Phonological Representation Learning for Isolated Signs Improves Out-of-Vocabulary Generalization

  • 引入语音学先验:参数解耦与半监督正则化
  • 在未见手势上实现更优的一次重建与识别效果
  • 适合关注手语模型泛化与语言结构建模的研究者

手语数据集词汇代表性不足,亟需具备泛化能力的模型。向量量化可学习离散的符号化表征,但其学习单元是否包含干扰泛化的虚假关联尚未明确。本文探究两种语音学归纳偏置:参数解耦(架构偏置)与语音学半监督(正则化技术),用于改进孤立手语识别与未见手势的重建质量。基于向量量化自编码器的实验表明,所提模型在已知手势识别上表现更优,且对未见手势的一次重建效果显著提升,表征更具判别性。该工作定量分析了显式语言学动机的归纳偏置如何提升手语表征的泛化能力。

原文摘要 · Abstract (English)

Sign language datasets are often not representative in terms of vocabulary, underscoring the need for models that generalize to unseen signs. Vector quantization is a promising approach for learning discrete, token-like representations, but it has not been evaluated whether the learned units capture spurious correlations that hinder out-of-vocabulary performance. This work investigates two phonological inductive biases: Parameter Disentanglement, an architectural bias, and Phonological Semi-Supervision, a regularization technique, to improve isolated sign recognition of known signs and reconstruction quality of unseen signs with a vector-quantized autoencoder. The primary finding is that the learned representations from the proposed model are more effective for one-shot reconstruction of unseen signs and more discriminative for sign identification compared to a controlled baseline. This work provides a quantitative analysis of how explicit, linguistically-motivated biases can improve the generalization of learned representations of sign language.

手语识别表征学习泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。