arXiv:2606.08056cs.CLcs.AI2026-06

解决手语中指代定位的识别难题,提升非词汇性表达建模能力。

What's the Point? Spatial Grammar & Index Resolution for Sign Language Processing

论文配图:What's the Point? Spatial Grammar & Index Resolution for Sign Language Processing
图 1 · 摘自论文原文
  • 将空间指代分解为检测与实体关联两步,实现精准定位。
  • 实测指代占手语10%-15%但现有模型恢复率极低。
  • 可作为插件增强已有手语识别模型,适用于结构化建模场景。

手语模型多依赖词元序列或文本监督,难以捕捉非词汇性、生成性表达。其中空间索引(如指向特定位置以标记话题实体)是典型挑战,而现有以词汇为中心的目标函数无法有效建模。本文针对手语识别中的索引现象开展专项评估,发现其虽占手语内容的10%-15%,但当前模型恢复效果不佳。为此提出一套训练与评估索引专家的框架,建立索引感知建模基线。方法将空间指代解析拆解为索引检测与话语实体链接两阶段,生成的提及表示支持自动标注和非词汇结构建模,并可作为辅助专家在推理时增强冻结的手语识别模型。

原文摘要 · Abstract (English)

Sign language models are predominantly trained with gloss-sequence or text supervision, thereby under-modeling non-lexical and productive constructions. One comparatively tractable instance is spatial indexing: pointing gestures that assign discourse entities to spatial loci for subsequent co-reference, which lexicon-centric objectives largely fail to capture. We present a targeted evaluation of indexing in Sign Language Recognition, showing that despite comprising 10-15% of signing content, indexing is poorly recovered. We introduce a framework for training and evaluating indexing experts, establishing a baseline for index-aware sign language modeling. Our approach decomposes spatial reference resolution into index detection and discourse entity linking. The resulting mention representations enable automatic annotation and non-lexical structure modeling, and serve as an auxiliary indexing expert that augments a frozen SLR model at inference time.

手语识别空间指代非词汇结构模型增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。