arXiv:2607.13372cs.CL2026-07

用小模型自动标注濒危语言,加词性标签能大幅减少人工工作量。

A POS Tier Is the Key to Automated Annotation for Low-Resource Language Documentation: Neural Interlinear Glossing for Irabu, a Southern Ryukyuan Language

论文配图:A POS Tier Is the Key to Automated Annotation for Low-Resource Language Documentation: Neural Interlinear Glossing for Irabu, a Southern Ryukyuan Language
图 1 · 摘自论文原文
  • 用小型BiLSTM-CRF模型实现分词、词性标注和释义的全流程自动标注。
  • 有词性标签时释义准确率提升4.4分,数据少时提升达11.6分。
  • 建议文档记录时四层标注:原文、词性、释义、翻译,效果最佳。

话语数据是田野语言学语法研究的主要实证基础,但生成逐行注释文本极为耗时,约需一小时处理一分钟录音。对濒危语言而言,与母语者验证分析的时间有限,自动化部分标注流程具有直接文献价值。本文针对琉球语南部方言伊鲁布语,构建了完整的神经标注流水线(分词、词性标注、释义),采用刻意小巧且透明的BiLSTM-CRF模型,并在严格限制下评估:仅使用约一小时完全标注的话语数据作为全部监督资源。实验操控两个变量:标注丰富度(是否含词性层)与数据量(训练预算从6到47分钟)。结果显示,有金标准词性标签时,语法释义准确率提升4.4分(标准差0.7),所有5个随机种子均显著;数据越少,增益越大(四分之一数据时提升11.6分);词性层使达到同等精度所需标注数据减半。但在全自动流程中该优势尚未实现:标注器在12%的词素上出错,错误词性比无词性更误导释义模型。该价值处于潜伏状态而非丢失:通过控制噪声模拟降级金标准发现,随着标注准确率提高,收益可恢复,当准确率达88%时出现拐点,92%-96%时可恢复1.6至3.2分。最终建议:语言记录应采用四层标注——文本、词性、释义、翻译。

原文摘要 · Abstract (English)

Discourse data are the primary empirical basis of grammar writing in field linguistics, but producing interlinearized text is notoriously expensive - on the order of one hour of work per minute of recording. For endangered languages, where the time remaining to verify analyses with native speakers is itself limited, automating parts of the interlinearization workflow has direct documentary value. We implement a full neural annotation pipeline (morpheme segmentation, POS tagging, glossing) for Irabu Ryukyuan using deliberately small, transparent BiLSTM-CRF models, and evaluate it under a realistic hard constraint: approximately one hour of fully annotated discourse as the entire supervised resource. Two factors of the annotation itself are manipulated: its richness (with or without a POS tier) and its quantity (training budgets from 6 to 47 minutes). Gold POS improves grammatical glossing by +4.4 (SD 0.7) points (significant in all 5 seeds), and the gain grows as data shrink (+11.6 points at a quarter of the data); a POS tier more than halves the amount of glossed data needed to reach a given accuracy. In a fully automatic pipeline this gain is not yet realized: the tagger still errs on 12% of morphemes, and an incorrect POS misleads the glossing model more than no POS at all. The value is latent rather than lost: degrading gold POS with controlled noise shows the gain returning as tagger accuracy rises, with break-even near our tagger's current 88% and +1.6 to +3.2 points recovered at 92-96%. We conclude with a concrete recommendation for documentation practice: annotate quadrilinearly - text, POS, gloss, translation.

低资源语言自动标注词性标注语言记录

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。