用专家纠错迭代训练,让古拉丁语语法分析器性能大幅提升
Correction as Annotation: Bootstrapping a Dependency Parser for Documentary Medieval Latin

- 用不靠谱模型预标注,专家纠错后反哺训练,循环迭代提升
- 句法准确率从48%升至92%,词性标注达98%,仅用1%数据量超越基线
- 适合古籍数字化、语言学研究者,尤其缺标注数据的冷门语种
中世纪文献的自然语言处理工具仍严重不足。现有五个可用的拉丁语树库模型在1258至1446年马赛编纂的160份文书上表现不佳,最佳标签依从率仅为0.62,形态感知得分仅0.24,性能与体裁或时期接近度无关。为弥补此差距,研究通过使用这些表现不佳的模型生成领域内训练数据:每轮迭代中,模型预标注200句,专家修正后用于训练下一模型,批次独立于模型状态采样,未采用主动学习选择。33小时标注投入覆盖1,804句,使通用词性标注准确率从0.80升至0.98,标签依从率从0.48升至0.92,所有指标均优于基线,且仅需最大基线模型1%的训练数据。标注人力占比从54%降至14-18%的平台期,该指标无需额外标准答案,可作停止依据。
原文摘要 · Abstract (English)
Medieval documentary sources remain inadequately served by existing natural language processing tools. None of the five readily available Latin treebank models attains usable performance on a collection of 160 inventories compiled in Marseille between 1258 and 1446. The best labelled attachment score is 0.62 and the best morphology-aware score is 0.24. Performance does not correlate with either genre or period proximity. To address this shortfall, in-domain training data was generated as a by-product of using these inadequate models. In each of nine iterations, a model pre-annotated 200 sentences; an expert corrected the annotations; and the corrected sentences were used to train the subsequent model, with batches sampled independently of model state, without active-learning selection. Thirty-three hours of annotation effort over 1,804 sentences increased universal part-of-speech accuracy from 0.80 to 0.98 and labelled attachment from 0.48 to 0.92, outperforming all baselines on the reported metrics while using 97% less training data than the largest one of them. Annotator effort declined from 54% of tokens to a plateau of 14-18%, an operational progress metric that requires no separate gold standard and can serve as a stopping criterion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。