arXiv:2606.28667cs.CL2026-06中稿 · CogSci 2026

探究手语识别模型是否真正理解音位特征,发现模型有音位感知但受架构限制。

Phonological Perception of Sign Language Models

论文配图:Phonological Perception of Sign Language Models
图 1 · 摘自论文原文
  • 用最小对立对测试模型对音位特征的敏感性。
  • 基于姿态的模型更关注手形差异,基于像素的模型更擅长捕捉位置变化。
  • 姿态模型表征与人类感知相似性相关(r~0.49),适合研究人机感知对比。

手语是组合系统,意义由手形、位置、动作等子音位参数组合生成。尽管手语识别(SLR)模型在翻译基准上表现提升,但其是否真正区分抽象音位特征,还是仅依赖低层统计相关性仍不明确。本研究通过最小对立对探针和与人类行为数据的表征对齐,评估了训练于美国手语(ASL)的SLR模型的音位感知能力。结果表明,SLR模型表现出涌现的音位敏感性,但存在明显架构权衡:基于姿态的模型对手形差异敏感,而基于像素的模型更善于捕捉位置变化。此外,姿态模型学习到的隐含表征与人类感知相似性判断呈显著相关(r~0.49)。这些发现表明,虽然当前SLR模型具备一定音位感知能力,但现有训练范式仍不足以突破其架构先验偏差的局限。

原文摘要 · Abstract (English)

Sign languages are compositional systems where meaning arises by combining sublexical phonological parameters, such as handshape, location, and movement. While deep learning models for Sign Language Recognition (SLR) have achieved increased performance on translation benchmarks, it remains unclear whether these models distinguish abstract phonological features or merely rely on low-level statistical correlations. This work evaluates the phonological perception of SLR models trained on American Sign Language (ASL) by probing phonological sensitivity using minimal pairs and evaluating representational alignment with human behavioral data. Our results reveal that SLR models exhibit emergent phonological sensitivity, but with clear architectural trade-offs: pose-based models are sensitive to handshape contrasts, while pixel-based models better capture location changes. Furthermore, pose-based models learn latent representations that correlate with human perceptual similarity judgments (r~0.49). These findings suggest that while SLR models exhibit emergent phonology, current training paradigms are insufficient to scale them beyond their architectural inductive biases.

手语识别音位感知深度学习表征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。