arXiv:2603.20114cs.CL2026-03

测试大模型对语法术语的翻译能力,发现准确率不足四成。

Current LLMs still cannot 'talk much' about grammar modules: Evidence from syntax

  • 用人类与ChatGPT-5对比翻译44个句法术语
  • 仅25%翻译准确,38.6%错误,36.4%部分正确
  • 建议语言学家与AI专家合作优化模型

我们旨在检验大型语言模型(LLMs)在描述语法模块方面的表达能力,基于ChatGPT将句法核心概念翻译为阿拉伯语的实证研究。从生成语法领域的书籍和期刊文章以及作者经验中收集了44个术语,由人工翻译后,再由ChatGPT-5进行翻译,随后采用分析与对比方法评估两组翻译结果。结果显示,大模型在处理嵌入术语中的核心句法属性时仍存在显著困难,仅有25%的翻译准确,38.6%不准确,36.4%部分正确,我们认为这部分可接受。基于此,提出一系列可操作策略,其中最突出的是促进人工智能专家与语言学家的紧密合作,以改进大模型对语法知识的处理机制,实现更准确或至少合理的翻译。

原文摘要 · Abstract (English)

We aim to examine the extent to which Large Language Models (LLMs) can 'talk much' about grammar modules, providing evidence from syntax core properties translated by ChatGPT into Arabic. We collected 44 terms from generative syntax previous works, including books and journal articles, as well as from our experience in the field. These terms were translated by humans, and then by ChatGPT-5. We then analyzed and compared both translations. We used an analytical and comparative approach in our analysis. Findings unveil that LLMs still cannot 'talk much' about the core syntax properties embedded in the terms under study involving several syntactic and semantic challenges: only 25% of ChatGPT translations were accurate, while 38.6% were inaccurate, and 36.4.% were partially correct, which we consider appropriate. Based on these findings, a set of actionable strategies were proposed, the most notable of which is a close collaboration between AI specialists and linguists to better LLMs' working mechanism for accurate or at least appropriate translation.

大模型句法翻译语言学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。