arXiv:2411.11260cs.CL2024-11被引 13

用大模型自动标注语法特征,90%准确率,省时省力。

Large corpora and large language models: a replicable method for automating grammatical annotation

  • 基于提示工程和小样本训练,用Claude 3.5 Sonnet自动标注语法结构
  • 在保留测试集上达到90%以上准确率,仅需少量标注数据
  • 适合大规模语料的语法研究,尤其适用于语言变异与演变分析

大量语言学研究依赖从文本语料中提取的标注数据,但语料规模的快速膨胀给语言学家的手动标注带来了实际困难。本文提出一种可复现的监督方法,借助大语言模型通过提示工程、训练与评估协助语法标注。以英语评价动词构式‘consider X (as) (to be) Y’的形式变异为案例,使用Claude 3.5 Sonnet模型及Davies的NOW与EnTenTen21(SketchEngine)语料数据。结果显示,在仅使用少量训练数据的情况下,模型在保留测试集上的准确率超过90%,验证了该方法未来大规模标注该构式的可行性。我们讨论了该方法对更广泛语法构式及语言变异与演变研究的泛化潜力,强调了人工智能协作者在语言学研究中的价值,尽管仍存在若干重要限制。

原文摘要 · Abstract (English)

Much linguistic research relies on annotated datasets of features extracted from text corpora, but the rapid quantitative growth of these corpora has created practical difficulties for linguists to manually annotate large data samples. In this paper, we present a replicable, supervised method that leverages large language models for assisting the linguist in grammatical annotation through prompt engineering, training, and evaluation. We introduce a methodological pipeline applied to the case study of formal variation in the English evaluative verb construction 'consider X (as) (to be) Y', based on the large language model Claude 3.5 Sonnet and corpus data from Davies' NOW and EnTenTen21 (SketchEngine). Overall, we reach a model accuracy of over 90% on our held-out test samples with only a small amount of training data, validating the method for the annotation of very large quantities of tokens of the construction in the future. We discuss the generalisability of our results for a wider range of case studies of grammatical constructions and grammatical variation and change, underlining the value of AI copilots as tools for future linguistic research, notwithstanding some important caveats.

大模型语法标注自然语言处理语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。