arXiv:2604.07189cs.CL2026-04被引 1

让大模型自主研究语料,自动发现语言演变规律。

Agent-Driven Corpus Linguistics: A Framework for Autonomous Linguistic Discovery

  • 大模型通过工具接口自主提出假设、查语料、分析结果
  • 在500万词语料中发现强度词演变链与语义变化路径
  • 适合语言学研究者快速获取可验证的语言发现

传统语料语言学依赖研究人员人工提出假设、构建查询并解释结果,耗时且需专业技能。本文提出代理驱动语料语言学框架:大语言模型(LLM)通过结构化工具接口连接语料查询引擎,自主完成假设生成、语料查询、结果解读与多轮迭代优化,人类研究者仅负责设定方向和评估最终输出。所有发现均基于可验证的语料证据,而非模型内生生成。该方法不替代理论与数据的关系定位,而是拓展了谁在进行研究这一维度。实验将LLM代理接入基于CQP索引的古腾堡语料库(500万词),仅输入“调查英语强化词”,即识别出历时传递链(so+ADJ > very > really)、三种语义变化路径(去词汇化、极性固化、隐喻约束)及语域敏感分布。对照实验表明,语料锚定带来量化与可证伪性,是训练数据无法提供的。为验证外部有效性,该代理在4000万词的CLMET语料库上复现了Claridge(2025)和De Smet(2013)两项研究,结果高度一致。该方法可在机器速度下产出实证发现,降低研究门槛。

原文摘要 · Abstract (English)

Corpus linguistics has traditionally relied on human researchers to formulate hypotheses, construct queries, and interpret results - a process demanding specialized technical skills and considerable time. We propose Agent-Driven Corpus Linguistics, an approach in which a large language model (LLM), connected to a corpus query engine via a structured tool-use interface, takes over the investigative cycle: generating hypotheses, querying the corpus, interpreting results, and refining analysis across multiple rounds. The human researcher sets direction and evaluates final output. Unlike unconstrained LLM generation, every finding is anchored in verifiable corpus evidence. We treat this not as a replacement for the corpus-based/corpus-driven distinction but as a complementary dimension: it concerns who conducts the inquiry, not the epistemological relationship between theory and data. We demonstrate the framework by linking an LLM agent to a CQP-indexed Gutenberg corpus (5 million tokens) via the Model Context Protocol (MCP). Given only "investigate English intensifiers," the agent identified a diachronic relay chain (so+ADJ > very > really), three pathways of semantic change (delexicalization, polarity fixation, metaphorical constraint), and register-sensitive distributions. A controlled baseline experiment shows that corpus grounding contributes quantification and falsifiability that the model cannot produce from training data alone. To test external validity, the agent replicated two published studies on the CLMET corpus (40 million tokens) - Claridge (2025) and De Smet (2013) - with close quantitative agreement. Agent-driven corpus research can thus produce empirically grounded findings at machine speed, lowering the technical barrier for a broader range of researchers.

语料语言学大模型自动发现语言演变

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。