arXiv:2510.24856cs.CL2025-10中稿 · publication in the…被引 1

用卢森堡语测试大模型语法理解,发现翻译好不等于懂语法。

Do Large Language Models Grasp The Grammar? Evidence from Grammar-Book-Guided Probing in Luxembourgish

  • 设计四阶段语法评测流程,以语法书为指导
  • 大模型在形态句法上表现弱,尤其错在最小对任务
  • 推理能力强的模型更可能掌握语法,适合语言研究者

语法是支配语言单位(如句子、短语、词语)结构组织与语义关系的规则体系。在自然语言处理中,针对语法的评估协议仍显匮乏,尤其在低资源语言中更为突出。大语言模型是否真正理解语法结构,特别是句法与语义之间的映射关系,仍存争议。为此,我们提出一种基于语法书的评估流程,构建系统化、可推广的语法评估框架,本研究以卢森堡语为例进行验证。结果表明,翻译性能与语法理解之间仅存在微弱正相关,说明强翻译能力并不等同于深层语法掌握。尽管大模型整体表现较好,主要得益于其语义优势,但在形态和句法层面依然薄弱,尤其在最小对任务中表现不佳;而较强的推理能力则为提升语法理解提供了可行路径。

原文摘要 · Abstract (English)

Grammar refers to the system of rules that governs the structural organization and the semantic relations among linguistic units such as sentences, phrases, and words within a given language. In natural language processing, there remains a notable scarcity of grammar focused evaluation protocols, a gap that is even more pronounced for low-resource languages. Moreover, the extent to which large language models genuinely comprehend grammatical structure, especially the mapping between syntactic structures and meanings, remains under debate. To investigate this issue, we propose a Grammar Book Guided evaluation pipeline intended to provide a systematic and generalizable framework for grammar evaluation consisting of four key stages, and in this work we take Luxembourgish as a case study. The results show a weak positive correlation between translation performance and grammatical understanding, indicating that strong translations do not necessarily imply deep grammatical competence. Larger models perform well overall due to their semantic strength but remain weak in morphology and syntax, struggling particularly with Minimal Pair tasks, while strong reasoning ability offers a promising way to enhance their grammatical understanding.

语法理解大模型评测低资源语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。