构建首个用于机器学习的古代伊比利亚语言数据集
Curation of a Palaeohispanic Dataset for Machine Learning
- 将散乱的古代伊比利亚语资料整理为机器学习可用的结构化数据
- 涵盖多种未完全破译的古伊比利亚语,支持跨语言分析
- 适合语言学与计算考古学研究者使用
古伊比利亚语言是公元前3世纪罗马人到来前在伊比利亚半岛使用的语言。其研究始于戈麦斯·莫诺发掘莱万特伊比利亚文字系统——一种半音节文字之一。然而,这些语言至今仍处于不同程度的未破译状态,无一完全可知。以往研究多基于纯语言学视角,计算方法尚未广泛应用。但现有资源匮乏且格式不适合机器学习技术。为此,本文构建了一个结构化数据集,旨在推动该领域的研究进展。
原文摘要 · Abstract (English)
Palaeohispanic languages are those spoken in the Iberian Peninsula before the arrival of the Romans in the 3rd Century B.C. Their study was really put on motion after Gómez Moreno deciphered the Iberian Levantine script, one of the several semi-sillabaries used by these languages. Still, the Palaeohispanic languages have varying degrees of decipherment, and none is fully known to this day. Most of the studies have been performed from a purely linguistic point of view, and a computational approach may benefit this research area greatly. However, the resources are limited and presented in an unsuitable format for techniques such as Machine Learning. Therefore, a structured dataset is constructed, which will hopefully allow more progress in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。