用大模型自动标注大规模语料,发现英语'consider'结构演变新规律。
A large-scale pipeline for LLM-assisted corpus annotation: variation and change in the English consider construction
- 四阶段流程:提示工程+预评估+批量处理+后验验证
- 60小时内完成14万条标注,准确率超98%
- 揭示不同文体中语法简化与强化的动态变化
随着自然语言语料规模空前扩张,人工标注仍是语料语言学研究的主要瓶颈。本文提出一种可扩展的大规模语言模型(LLM)辅助语料标注流水线,通过四阶段工作流——提示工程、事前评估、自动化批处理和事后验证——实现海量语料的语法标注。以美国历史语料库(COHA)中的143,933条'consider'例句为对象,利用OpenAI API在60小时内完成标注,对两种复杂标注任务的准确率均达98%以上。基于44,527个真实阳性样本构建贝叶斯多项式GAM模型,揭示了评价性'consider X as/to be/Ø Y'结构在不同文体中的演化轨迹,支持关于语域正式程度与形态句法简化/增强竞争关系的新假说。结果表明,大模型可在极小人力干预下完成规模化数据准备,使以往难以实现的研究问题成为可能,但需关注成本、许可与伦理问题。
原文摘要 · Abstract (English)
As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical annotation in voluminous corpora using large language models (LLMs). Unlike previous supervised and iterative approaches, our method employs a four-phase workflow: prompt engineering, pre-hoc evaluation, automated batch processing, and post-hoc validation. We demonstrate the pipeline's accessibility and effectiveness through a diachronic case study of variation in the English evaluative consider construction (consider X as/to be/Ø Y). We annotate 143,933 'consider' concordance lines from the Corpus of Historical American English (COHA) via the OpenAI API in under 60 hours, achieving 98%+ accuracy on two sophisticated annotation procedures. A Bayesian multinomial GAM fitted to 44,527 true positives of the evaluative construction reveals previously undocumented genre-specific trajectories of change, enabling us to advance new hypotheses about the relationship between register formality and competing pressures of morphosyntactic reduction and enhancement. Our results suggest that LLMs can perform a range of data preparation tasks at scale with minimal human intervention, unlocking substantive research questions previously beyond practical reach, though implementation requires attention to costs, licensing, and other ethical considerations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。