arXiv:2511.00343cs.CL2025-11EMNLP被引 5

测试大模型能否像语言学家一样推理语法结构

LingGym: How Far Are LLMs from Thinking Like Field Linguists?

  • 用跨语言语法数据评估模型的元语言推理能力
  • 加入结构化语言线索后性能显著提升
  • 适合关注低资源语言研究的学者参考

本文提出LingGym,一个新基准,通过18种类型多样的参考语法中提取的逐行注释文本(IGT)和语法描述,评估大语言模型在元语言推理方面的能力。与以往聚焦特定下游任务的研究不同,本工作考察模型是否能在未见的低资源语言及结构上进行泛化推理。我们设计了受控评估任务:词-注释推断,要求模型根据上下文推断缺失的词汇及其注释,使用不同层级的语言信息(如注释、语法解释、翻译)。结果表明,引入结构化语言线索可显著提升所有模型的推理表现。该研究揭示了大语言模型在类型学导向的语言分析与低资源语言记录中的潜力与当前局限。

原文摘要 · Abstract (English)

This paper introduces LingGym, a new benchmark that evaluates LLMs' capacity for meta-linguistic reasoning using Interlinear Glossed Text (IGT) and grammatical descriptions extracted from 18 typologically diverse reference grammars. Unlike previous work that focuses on specific downstream tasks, we assess whether LLMs can generalize linguistic inference across low-resource languages and structures not seen during training. We present a controlled evaluation task: Word-Gloss Inference, in which the model must infer a missing word and gloss from context using varying levels of linguistic information (e.g., glosses, grammatical explanations, translations). Our results show that incorporating structured linguistic cues leads to consistent improvements in reasoning performance across all models. This work highlights both the promise and current limitations of using LLMs for typologically informed linguistic analysis and low-resource language documentation.

语言模型元语言推理低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。