用统一音素嵌入解决跨文字地名匹配难题,无需语言识别或音标资源。
Symphonym: Universal Phonetic Embeddings for Cross-Script Name Matching
- 通过师生知识蒸馏,将20种文字地名映射到128维音素空间。
- 在中世纪希伯来与阿拉伯地名测试集上达85.2%召回率和90.8%平均倒数排名。
- 可处理历史文献拼写变异,适用于档案中人名与地名解析。
跨书写系统地名匹配是整合多语言地理数据的长期挑战,涵盖现代地名录、中世纪旅行记录及殖民时期调查。现有方法依赖语言特定音标算法或罗马化步骤,会丢失音素信息,且无法跨文字边界泛化。本文提出Symphonym,一种神经嵌入系统,将20种文字的地名映射至统一的128维音素空间,实现无需语言识别或音标资源的直接跨文字相似性比较。采用教师-学生知识蒸馏架构,先从IPA转录的发音特征学习,再将知识迁移至字符级学生模型。在包含6700万条地名的GeoNames、Wikidata与Getty Thesaurus of Geographic Names数据集上,使用3270万组三元组样本训练,学生模型在MEHDIE跨文字基准测试(由领域专家构建的中世纪希伯来与阿拉伯地名匹配集)上取得最高召回率@1(85.2%)和平均倒数排名(90.8%),证明了从现代训练材料向史前资料的跨时间泛化能力。仅用原始发音特征的消融实验仅达45.0% MRR,证实神经训练课程的有效性。该方法自然处理历史文档中的非标准化拼写差异,并成功迁移至档案中的个人姓名,表明其在数字人文与开放链接数据场景下具广泛适用性。
原文摘要 · Abstract (English)
Matching place names across writing systems is a persistent obstacle to the integration of multilingual geographic sources, whether modern gazetteers, medieval itineraries, or colonial-era surveys. Existing approaches depend on language-specific phonetic algorithms or romanisation steps that discard phonetic information, and none generalises across script boundaries. This paper presents Symphonym, a neural embedding system which maps toponyms from twenty writing systems into a unified 128-dimensional phonetic space, enabling direct cross-script similarity comparison without language identification or phonetic resources at inference time. A Teacher-Student knowledge distillation architecture first learns from articulatory phonetic features derived from IPA transcriptions, then transfers this knowledge to a character-level Student model. Trained on 32.7 million triplet samples drawn from 67 million toponyms spanning GeoNames, Wikidata, and the Getty Thesaurus of Geographic Names, the Student achieves the highest Recall@1 (85.2%) and Mean Reciprocal Rank (90.8%) on the MEHDIE cross-script benchmark -- medieval Hebrew and Arabic toponym matches curated by domain experts and entirely independent of the training data -- demonstrating cross-temporal generalisation from modern training material to pre-modern sources. An ablation using raw articulatory features alone yields only 45.0% MRR, confirming the contribution of the neural training curriculum. The approach naturally handles pre-standardisation orthographic variation characteristic of historical documents, and transfers effectively to personal names in archival sources, suggesting broad applicability to name resolution tasks in digital humanities and linked open data contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。