GlossAssist让语言学家能快速标注文本并持续改进AI模型。
GlossAssist -- A Tool to Simplify Corpus Creation and Study the Effect of NLP Models in Low-Resource Documentation Settings

- 基于CWoMP的检索架构,用可编辑词素库生成注释。
- 标注修改直接更新词素库,无需重训练提升后续效果。
- 专为田野语言学家设计,支持人机协作迭代优化。
互文标注文本(IGT)是语言记录的标准标注格式。然而,手动制作耗时且成本高。尽管自动化标注系统近年已有显著进步,但田野语言学家使用率仍低。现有工具多用于评估而非实际使用,缺乏可解释的修正路径,也无法将语言学知识反馈至模型行为。本文提出GlossAssist,一个基于CWoMP(对比词素预训练)检索架构的标注工具,其预测基于可编辑的词素表示词库。结合CWoMP,系统将标注者的每一次修正视为主动学习的一部分,动态扩展词库以改进未来预测,无需重新训练模型。本文介绍该界面,并主张这种反馈循环应作为面向语言记录者的NLP工具的设计核心。
原文摘要 · Abstract (English)
Interlinear glossed text (IGT) is the standard format for linguistic annotation in language documentation. Producing it manually, however, is often slow and costly. Automated glossing systems have improved substantially in recent years, but adoption among field linguists remains limited. Existing tools are designed to be evaluated rather than used, offering no interpretable path for correction or the incorporation of linguistic expertise back into model behavior. We present GlossAssist, a glossing tool built around the retrieval-based architecture of CWoMP (Contrastive Word-Morpheme Pre-training), which grounds predictions in a mutable lexicon of learned morpheme representations. In conjunction with CWoMP, our system treats each correction by an annotator as part of an active learning setting, which expands the lexicon and improves future predictions without having to retrain the model. In this paper, we present our interface and argue that this feedback loop should be treated as a design requirement for NLP tools aimed at documentary linguists.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。