arXiv:2603.00923cs.CL2026-03被引 2

用混合模型自动标注濒危语言形态,大幅减少人工工作量。

Hybrid Neural-LLM Pipeline for Morphological Glossing in Endangered Language Documentation: A Case Study of Jungar Tuvan

  • 先用神经网络初标,再用大模型纠错,两阶段协同提升精度。
  • 少样本示例越多,效果越好,但增益呈对数递减趋势。
  • 反直觉发现:词素词典反而降低性能,适合低资源语言标注场景。

互线标注文本(IGT)的生成仍是语言学记录与田野调查中的主要瓶颈,尤其对低资源、形态复杂的语言而言。本文提出一种混合自动标注流水线,结合神经序列标注与大语言模型(LLM)后处理修正,以低资源突厥语种——准格尔图瓦语为案例进行评估。系统性消融实验表明,检索增强提示相比随机示例选择显著提升效果;更意外的是,在多数情况下,提供词素词典反而比不提供更差;性能随少量示例数量近似对数增长。最显著的是,基于BiLSTM-CRF模型与LLM后修正的两阶段流程在多数模型中均取得显著提升,有效降低标注工作量。基于此,我们提炼出在形态复杂语言田野工作中融合结构化预测模型与LLM推理的设计原则。这些原则表明,混合架构是实现计算轻量化的自动语言标注的可行路径,尤其适用于濒危语言记录。

原文摘要 · Abstract (English)

Interlinear glossed text (IGT) creation remains a major bottleneck in linguistic documentation and fieldwork, particularly for low-resource morphologically rich languages. We present a hybrid automatic glossing pipeline that combines neural sequence labeling with large language model (LLM) post-correction, evaluated on Jungar Tuvan, a low-resource Turkic language. Through systematic ablation studies, we show that retrieval-augmented prompting provides substantial gains over random example selection. We further find that morpheme dictionaries paradoxically hurt performance compared to providing no dictionary at all in most cases, and that performance scales approximately logarithmically with the number of few-shot examples. Most significantly, our two-stage pipeline combining a BiLSTM-CRF model with LLM post-correction yields substantial gains for most models, achieving meaningful reductions in annotation workload. Drawing on these findings, we establish concrete design principles for integrating structured prediction models with LLM reasoning in morphologically complex fieldwork contexts. These principles demonstrate that hybrid architectures offer a promising direction for computationally light solutions to automatic linguistic annotation in endangered language documentation.

形态标注低资源混合模型濒危语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。