构建德语与德国手语词义对应数据集,揭示多对一等映射规律。
How Do Lexical Senses Correspond Between Spoken German and German Sign Language?
- 人工标注1404组词义-手语映射,识别三类对应模式。
- 语义相似度方法准确率达88.52%,显著优于精确匹配。
- 首个跨模态词义对应数据集,适合语言学与AI研究者使用。
手语词典编纂依赖词与手语符号的映射,但多义词在不同语境下对应不同手语的现象常被忽略。本文通过分析德语与德国手语(DGS),从德语用法图谱(D-WUG)中选取32个词、数字手语词典(DW-DGS)中选取49个手语,人工标注1404组词义-手语映射,识别出三类对应关系:类型1(一词对多手语)、类型2(多词对一手语)、类型3(一对一)及无匹配情况。评估了精确匹配(EM)和基于SBERT嵌入的语义相似度(SS)两种计算方法,结果显示SS整体准确率88.52%远超EM的71.31%,尤其在类型1上提升52.1个百分点。本研究首次建立跨模态词义对应标注数据集,揭示哪些映射模式可被计算识别,代码与数据已公开。
原文摘要 · Abstract (English)
Sign language lexicographers construct bilingual dictionaries by establishing word-to-sign mappings, where polysemous and homonymous words corresponding to different signs across contexts are often underrepresented. A usage-based approach examining how word senses map to signs can identify such novel mappings absent from current dictionaries, enriching lexicographic resources. We address this by analyzing German and German Sign Language (Deutsche Gebärdensprache, DGS), manually annotating 1,404 word use-to-sign ID mappings derived from 32 words from the German Word Usage Graph (D-WUG) and 49 signs from the Digital Dictionary of German Sign Language (DW-DGS). We identify three correspondence types: Type 1 (one-to-many), Type 2 (many-to-one), and Type 3 (one-to-one), plus No Match cases. We evaluate computational methods: Exact Match (EM) and Semantic Similarity (SS) using SBERT embeddings. SS substantially outperforms EM overall 88.52% vs. 71.31%), with dramatic gains for Type 1 (+52.1 pp). Our work establishes the first annotated dataset for cross-modal sense correspondence and reveals which correspondence patterns are computationally identifiable. Our code and dataset are made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。