用神经模型辅助语言田野调查,提升形态数据收集效率
Can a Neural Model Guide Fieldwork? A Case Study on Morphological Data Collection
- 基于范式表单元的均匀采样提升数据多样性
- 利用模型置信度引导标注,增强研究者与说话人互动
- 适合语言学田野工作者与计算语言学研究者
语言田野调查是语言记录与保护的重要环节,但过程漫长且耗时。本文提出一种新型模型,用于指导语言学家在田野工作中收集形态数据,并考虑研究者与说话人之间的动态交互。我们构建了一个新框架,评估不同采样策略在获取形态数据方面的效率,并检验当前先进神经模型对形态结构的泛化能力。实验表明,提高效率的两个关键策略为:(1) 在范式表的各个单元中采用均匀采样以增加标注数据多样性;(2) 利用模型置信度作为指引,在标注过程中提供可靠预测,从而增强正向互动。
原文摘要 · Abstract (English)
Linguistic fieldwork is an important component in language documentation and preservation. However, it is a long, exhaustive, and time-consuming process. This paper presents a novel model that guides a linguist during the fieldwork and accounts for the dynamics of linguist-speaker interactions. We introduce a novel framework that evaluates the efficiency of various sampling strategies for obtaining morphological data and assesses the effectiveness of state-of-the-art neural models in generalising morphological structures. Our experiments highlight two key strategies for improving the efficiency: (1) increasing the diversity of annotated data by uniform sampling among the cells of the paradigm tables, and (2) using model confidence as a guide to enhance positive interaction by providing reliable predictions during annotation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。