arXiv:2502.00090cs.CL2025-02被引 2

破解古埃兰文数字歧义,助力古代账目文献解读

Disambiguating Numeral Sequences to Decipher Ancient Accounting Corpora

  • 基于文档结构与自举学习构建两种歧义消解方法
  • 发现泥板内容与数字大小存在未知关联
  • 为古代会计文献研究提供关键工具和测试集

记数系统将抽象数值编码为具体的文字序列。现代书写系统的记数体系通常精确无歧义,但古部分破译的原始埃兰文(PE)记数系统中,同一数字符号可能有最多四种不同读法,取决于解读方式。本文旨在对这些读法进行歧义消解,以确定该文献中记录的数值。我们算法提取每个PE数字标记的所有可能读法,并提出两种基于原始文档结构特征及自举训练分类器的消歧技术。此外,我们构建了一个用于评估消歧方法的测试集,以及一种谨慎选择自举分类器规则的新方法。分析验证了现有对这一文字系统的直觉,并揭示了泥板内容与数字大小间此前未知的相关性。该工作对理解与破译原始埃兰文至关重要,因其文献以会计为主,数字符号数量远超文本符号。

原文摘要 · Abstract (English)

A numeration system encodes abstract numeric quantities as concrete strings of written characters. The numeration systems used by modern scripts tend to be precise and unambiguous, but this was not so for the ancient and partially-deciphered proto-Elamite (PE) script, where written numerals can have up to four distinct readings depending on the system that is used to read them. We consider the task of disambiguating between these readings in order to determine the values of the numeric quantities recorded in this corpus. We algorithmically extract a list of possible readings for each PE numeral notation, and contribute two disambiguation techniques based on structural properties of the original documents and classifiers learned with the bootstrapping algorithm. We also contribute a test set for evaluating disambiguation techniques, as well as a novel approach to cautious rule selection for bootstrapped classifiers. Our analysis confirms existing intuitions about this script and reveals previously-unknown correlations between tablet content and numeral magnitude. This work is crucial to understanding and deciphering PE, as the corpus is heavily accounting-focused and contains many more numeric tokens than tokens of text.

古文字破译数字歧义会计文献自举学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。