arXiv:2608.01935cs.CLcs.AI2026-08

首个通用古希腊语元音长短自动标注工具,解决文本语音标注难题。

Automatic Annotation of Ancient Greek Vowel Length

  • 基于词干、词性等信息,通过递归模块继承高频词标注,实现跨词形推广。
  • 在诗体与散文基准上,生成数据训练的Transformer模型准确率超过规则系统。
  • 适合古希腊语语音学、文体分析及韵律建模研究者使用。

古希腊语自然语言处理依赖的语料库通常无法区分字母α、ι、υ的音位长短,统称dichrona。这些字母在不同词义、形态、音变、句法及时期、体裁、诗体惯例下可表示长音或短音。正确判断并标记其长度称为“macronizing”,因词形繁多且高度依赖上下文,属长尾难题。目前尚无大规模公开的macronized古希腊语语料库,亟需独立的macronizer。此前工作仅能构建特定语料库的静态标注词典,本文首次构建了适用于任意古希腊语输入的通用macronizer。输入包含词干、词性及形态标注(标准CoNLL-U格式),通过一组递归模块,使罕见词形可继承同词根高频词的标注。该macronizer主要用途是生成机器学习训练数据:我们展示,仅用其输出训练的小型字符级Transformer,能泛化至规则系统未标注的案例,在诗体与散文的金标准人工标注基准上,准确率匹配甚至超越原规则系统。同时证明,元音长短标注可提升下游韵律任务(如诗歌节奏分析)的性能。

原文摘要 · Abstract (English)

Prior work in Ancient Greek NLP relies on corpora that do not disambiguate the phonemic vowel length of alpha, iota, and ypsilon, together known as the dichrona. Depending on lexeme, morphology, sandhi, syntax, and conventions of period, genre, and verse form, each of these letters can represent either a long or a short vowel. Deciding and marking the correct length is known as "macronizing", a long-tail problem given the sheer mass of word forms and the context dependency of individual instances. No macronized corpus of Ancient Greek is publicly available at scale, so a stand-alone macronizer is needed. While previous work has shown how to build a static, corpus-bespoke vowel-length dictionary, the present paper constructs the first general-purpose macronizer for arbitrary Ancient Greek input. Given input carrying lemma, part-of-speech, and morphological annotation in the standard CoNLL-U format, a set of recursive modules lets less common word forms inherit markup from more common forms of the same lexical word. The macronizer's chief application is generating training data for machine learning: we show that a small character-level transformer trained on the macronizer's own output learns to generalize past the cases the rule-based system leaves unmarked, matching or exceeding its accuracy on a gold-standard, manually annotated benchmark of verse and prose. We also show that macronization can improve downstream prosodical NLP tasks like verse scansion.

古希腊语语音标注规则系统韵律分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。