arXiv:2609.07791cs.CL2026-09

用大模型代理自动分析语言类型学特征,提升跨语言研究效率。

LLM Agents as Computational Typologists

论文配图:LLM Agents as Computational Typologists
图 1 · 摘自论文原文
  • 构建基于ReAct框架的智能代理,可检索语料、解析句式并推理假设。
  • 在25个开源语法书中验证,能识别支持与反例,准确率超人工基准。
  • 适合语言学家快速生成假设,但需专家复核关键结论。

语言类型学依赖专家对多种语言参考语法的分析,导致大规模跨语言比较耗时且难以扩展。我们提出AUTOTYPOLOGIST,一种基于大模型的代理系统,可在参考语法中进行证据驱动的类型学分析。该代理能检索相关语法段落,分析逐行释义文本(IGT),并通过类似ReAct的工作流迭代推理类型学假设。我们在25个开源参考语法上评估了该系统在类型特征标注(TYPOLOGICAL FEATURE CODING)和类型普遍性检验(TYPOLOGICAL HYPOTHESIS TESTING)上的表现。在信息受限条件下,代理虽能整合语法描述信息,但在仅提供目标语言IGT时仍面临挑战;在跨语言证据合成方面,代理能有效识别支持案例与反例。结果表明,大模型代理可支持可扩展且可追溯的类型学分析,但仍需专家验证。

原文摘要 · Abstract (English)

Linguistic typology relies on expert analysis of reference grammars across languages, making large-scale crosslinguistic comparison labor-intensive and unscalable. We introduce AUTOTYPOLOGIST, an LLM agent for evidence-grounded typological analysis over reference grammars. The agent is capable of retrieving relevant grammar sections, analyzing interlinear glossed text (IGT), and iteratively reasoning over typological hypotheses using a ReAct-style workflow. We evaluate the system on TYPOLOGICAL FEATURE CODING against expert annotations and TYPOLOGICAL HYPOTHESIS TESTING with typological universals using 25 open-source reference grammars. Operating under different information constraints in TYPOLOGICAL FEATURE CODING, the agent can synthesize information from reference grammar prose but still faces challenges with only IGTs in the target language. In TYPOLOGICAL HYPOTHESIS TESTING, the agent can synthesize crosslinguistic evidence and identify both supporting cases and counterexamples. These findings suggest that LLM agents can support scalable and inspectable typological analysis, while still requiring expert validation.

语言学LLM代理类型学自动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。