arXiv:2503.19260cs.CLcs.AI2025-03被引 6

大模型在语法标注任务中存在明显盲区,难以准确识别复杂句法结构。

Linguistic Blind Spots of Large Language Models

  • 通过实验检验大模型对词性、短语和从句的识别能力
  • 最强模型Llama3-70b仍会误判嵌套从句和动词短语
  • 适合关注大模型语言理解局限的研究者参考

大型语言模型(LLMs)是当前许多AI应用的基础。尽管它们在生成连贯文本方面表现出色,但其在细粒度语言标注任务中的表现仍存疑,如名词或动词的检测,以及输入文本中复杂句法结构(如从句)的识别。这些任务需要精确的句法与语义理解。当大模型在特定语言结构上表现不佳时,其输出可靠性令人质疑。本文通过一系列实验,研究了近期大模型在细粒度语言标注任务中的表现。结果表明,大模型在处理语言复杂输入时效率有限,最强大的模型(Llama3-70b)仍会出现显著错误,例如误判嵌套从句、无法识别动词短语,或将复杂名词短语混淆为从句。这些发现为未来大模型设计提供了重要启示。

原文摘要 · Abstract (English)

Large language models (LLMs) are the foundation of many AI applications today. However, despite their remarkable proficiency in generating coherent text, questions linger regarding their ability to perform fine-grained linguistic annotation tasks, such as detecting nouns or verbs, or identifying more complex syntactic structures like clauses in input texts. These tasks require precise syntactic and semantic understanding of input text, and when LLMs underperform on specific linguistic structures, it raises concerns about their reliability for detailed linguistic analysis and whether their (even correct) outputs truly reflect an understanding of the inputs. In this paper, we empirically study the performance of recent LLMs on fine-grained linguistic annotation tasks. Through a series of experiments, we find that recent LLMs show limited efficacy in addressing linguistic queries and often struggle with linguistically complex inputs. We show that the most capable LLM (Llama3-70b) makes notable errors in detecting linguistic structures, such as misidentifying embedded clauses, failing to recognize verb phrases, and confusing complex nominals with clauses. Our results provide insights to inform future advancements in LLM design and development.

语言模型句法分析大模型缺陷语言理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。