arXiv:2502.10934cs.CL2025-02被引 7

o3模型无法理解语言深层结构,暴露了当前AI在组合性推理上的根本缺陷。

Fundamental Principles of Linguistic Structure are Not Represented by o3

  • 测试发现o3依赖表层统计,无法掌握句法结构规则
  • 在语义非法的比较句中表现失败,无法正确判断可接受性
  • 无法区分语法与语义错误,缺乏多解析评估能力

具备快速构建和操作具身组合抽象的能力,以及在人类语言创造性使用所需的递归层级句法对象上展现专业性的能力,是实现通用人工智能的核心要素。我们评估了近期发布的o3模型(OpenAI;o3-mini-high),发现尽管它在部分依赖线性表面统计的简单语言测试(如草莓测试)中表现良好,但在泛化基本短语结构规则方面失败;在涉及语义上非法基数比较的‘埃舍尔句’中失败;无法正确评估并解释可接受性动态;也无法区分生成语义不合法与句法不合法输出的指令。当被要求生成违反语法规则的简单例子时,它似乎无法表示多个解析以供不同语义解释的评估。与近年来声称语言模型即将取代语言学领域的观点形成鲜明对比,我们的结果表明,深度学习不仅在组合性方面遭遇瓶颈(Marcus 2022),更撞上了一个顽固且难以突破的壁垒,仅靠更多算力无法轻易跨越,达到类人组合推理水平。

原文摘要 · Abstract (English)

A core component of a successful artificial general intelligence would be the rapid creation and manipulation of grounded compositional abstractions and the demonstration of expertise in the family of recursive hierarchical syntactic objects necessary for the creative use of human language. We evaluated the recently released o3 model (OpenAI; o3-mini-high) and discovered that while it succeeds on some basic linguistic tests relying on linear, surface statistics (e.g., the Strawberry Test), it fails to generalize basic phrase structure rules; it fails with comparative sentences involving semantically illegal cardinality comparisons ('Escher sentences'); its fails to correctly rate and explain acceptability dynamics; and it fails to distinguish between instructions to generate unacceptable semantic vs. unacceptable syntactic outputs. When tasked with generating simple violations of grammatical rules, it is seemingly incapable of representing multiple parses to evaluate against various possible semantic interpretations. In stark contrast to many recent claims that artificial language models are on the verge of replacing the field of linguistics, our results suggest not only that deep learning is hitting a wall with respect to compositionality (Marcus 2022), but that it is hitting [a [stubbornly [resilient wall]]] that cannot readily be surmounted to reach human-like compositional reasoning simply through more compute.

语言模型组合性句法理解AI局限

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。