用课堂数据构建古英语树库,验证新手标注与LLM后编辑的有效性。
Building UD Cairo for Old English in the Classroom
- 结合LLM提示与真实文本收集20句古英语例句。
- 新手标注经整合后可得高质量结果,且能提升语言认知。
- 用现代英语训练数据解析古英语特征,性能有提升。
本文基于历史语言学课程中的课堂实践,构建了一个以UD Cairo句子为基础的古英语样本树库。通过结合大模型提示与真实古英语语料搜索,选取了20个涵盖多种句法结构的句子进行收集。标注由多名对UD接触有限的学生完成,其结果经比对和裁决。结果显示,当前大模型生成的古英语语法不反映真实句法,但可通过后编辑改善;尽管初学者无法完美完成标注任务,但整体协作可产出良好成果并从中学习。此外,我们使用现代英语训练数据进行了初步解析实验,发现直接用于古英语时性能较差,但若解析词形(lemma)、超词形(hyperlemma)及释义(gloss)等标注特征,则性能有所提升。
原文摘要 · Abstract (English)
In this paper we present a sample treebank for Old English based on the UD Cairo sentences, collected and annotated as part of a classroom curriculum in Historical Linguistics. To collect the data, a sample of 20 sentences illustrating a range of syntactic constructions in the world's languages, we employ a combination of LLM prompting and searches in authentic Old English data. For annotation we assigned sentences to multiple students with limited prior exposure to UD, whose annotations we compare and adjudicate. Our results suggest that while current LLM outputs in Old English do not reflect authentic syntax, this can be mitigated by post-editing, and that although beginner annotators do not possess enough background to complete the task perfectly, taken together they can produce good results and learn from the experience. We also conduct preliminary parsing experiments using Modern English training data, and find that although performance on Old English is poor, parsing on annotated features (lemma, hyperlemma, gloss) leads to improved performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。