用广播新闻文本弱监督训练手语表示,跨数据集定位效果显著提升。
Gloss-Free Representation Learning for Cross-Dataset Sign Spotting

- 用语音转录文本生成伪词标签,无需人工标注
- 在土耳其手语数据上,定位准确率(IoU)提升至0.465
- 适合资源匮乏语言的手语表示学习,尤其适用于形态丰富的语言
资源受限的手语研究常受限于密集语言标注(如词义、时间边界)的成本。广播新闻提供了一种替代方案:连续手语与口语转录对齐,但这种监督较弱,因文本与手语仅松散对应。对于形态丰富的语言如土耳其语,同一词汇可有多种变体,而部分派生形式需保持区分。本文研究是否可利用弱文本监督预训练可复用的手语编码器。基于新构建的土耳其广播语料库TSL-News,采用转录文本生成伪词标签,对比规则化词干提取与基于约束大模型的归一化方法。在新构建的TSL Spotting Benchmark上评估,使用大模型辅助归一化的编码器使顶5定位平均交并比(mean IoU)从0.235提升至0.465,56.2%样本达到IoU≥0.5。频率分析表明该提升非源于记忆高频伪词。下游翻译任务中,BLEU-4由9.60升至11.04,ROUGE由23.48升至27.43。结果表明,松散对齐的广播数据可有效提供弱监督,用于学习捕捉词汇内容与时间结构的手语表示。
原文摘要 · Abstract (English)
Sign-language research for resource-constrained languages is often limited by the cost of dense linguistic labels such as glosses, temporal boundaries, and sign order. Broadcast news offers a practical alternative by pairing continuous signing with spoken-language transcripts, but this supervision is weak since text and signing are loosely aligned. Morphologically rich languages such as Turkish add further difficulty, as the same lexical meaning can appear in many inflected forms while some derived forms should remain distinct. We study whether weak transcript-based supervision can pretrain a reusable sign encoder in this setting, where poor text normalization can fragment pseudo-gloss targets and weaken representation learning. Unlike prior pseudo-gloss pipelines designed mainly to improve translation, we test whether the pretrained encoder transfers as a reusable representation for cross-dataset sign spotting. We pretrain on TSL-News, a new Turkish broadcast corpus, using pseudo-gloss labels derived from transcripts rather than manual annotation, comparing rule-based morphological lemmatization with constrained LLM-assisted normalization over a fixed vocabulary. We evaluate the learned representations via cross-dataset sign spotting on a new TSL Spotting Benchmark built from the TSL Dictionary corpus. The LLM-assisted encoder raises top-5 temporal localization mean IoU from 0.235 to 0.465, with 56.2% of examples reaching an IoU of at least 0.50; a frequency analysis suggests this gain is not mainly driven by memorizing frequent pseudo-gloss labels. In a downstream translation check, the same pretraining improves BLEU-4 from 9.60 to 11.04 and ROUGE from 23.48 to 27.43. These results show that loosely aligned broadcast data can provide effective weak supervision for learning sign representations that capture both lexical content and temporal structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。