用智能大模型自动分析多语言语法结构,让语言研究更高效。
Towards Corpus-Grounded Agentic LLMs for Multilingual Grammatical Analysis
- 构建可理解指令、生成代码并基于语料推理的智能框架
- 在170+语言上测试,准确识别13类词序特征的主次顺序
- 适合语言学研究者与跨语言数据自动化分析场景
经验性语法研究日益依赖数据,但对标注语料的系统分析仍需大量方法与技术投入。本文探索如何利用智能型大语言模型(LLMs)通过推理标注语料,回答语言学问题并生成可解释、数据驱动的答案。提出一个基于语料的语法分析智能框架,融合自然语言任务解析、代码生成与数据驱动推理等概念。以通用依存关系(Universal Dependencies, UD)语料为实验基础,针对世界语言结构图谱(WALS)启发的多语言语法任务进行验证。评估涵盖13种词序特征和超过170种语言,从三个维度衡量系统表现:主导词序准确率、词序覆盖完整性和分布保真度,反映系统泛化能力、识别能力和量化差异的能力。结果表明,结合大模型推理与结构化语言数据具有可行性,为可解释、可扩展的语料驱动语法研究提供了初步路径。
原文摘要 · Abstract (English)
Empirical grammar research has become increasingly data-driven, but the systematic analysis of annotated corpora still requires substantial methodological and technical effort. We explore how agentic large language models (LLMs) can streamline this process by reasoning over annotated corpora and producing interpretable, data-grounded answers to linguistic questions. We introduce an agentic framework for corpus-grounded grammatical analysis that integrates concepts such as natural-language task interpretation, code generation, and data-driven reasoning. As a proof of concept, we apply it to Universal Dependencies (UD) corpora, testing it on multilingual grammatical tasks inspired by the World Atlas of Language Structures (WALS). The evaluation spans 13 word-order features and over 170 languages, assessing system performance across three complementary dimensions - dominant-order accuracy, order-coverage completeness, and distributional fidelity - which reflect how well the system generalizes, identifies, and quantifies word-order variations. The results demonstrate the feasibility of combining LLM reasoning with structured linguistic data, offering a first step toward interpretable, scalable automation of corpus-based grammatical inquiry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。