数据驱动模型无法实现符号级逻辑推理,因训练数据不足且目标矛盾。
Data-driven Machine Learning Cannot Reach Symbolic-level Logical Reasoning -- The Limit of the Scaling Law

- 用球面向量显式构建模型,突破传统数据依赖。
- GPT-5和Euler Net均在100%准确率时出现幻觉,无法深化理解。
- 即使训练集无限扩大,也无法保证符号级推理能力。
通过将向量映射到球面并实现显式模型构建,神经网络可在无训练数据的情况下完成符号级三段论推理。我们发现,传统数据驱动机器学习系统存在两大根本局限:组合表生成的训练数据无法区分全部24种有效三段论类型;端到端前提到结论的映射在神经组件内产生矛盾目标。对两种代表性系统(使用语言输入的GPT-5和使用视觉输入的Euler Net)的实验支持此分析。ChatGPT GPT-5虽可达100%准确率,但伴随幻觉现象;因学习过程在100%准确时终止,系统无法从经验性准确跃迁至符号级推理。随机测试数据使Euler Net准确率降至56%;反复扩增训练集可将其提升至97%,并在8种三段论类型上表现完美。然而,由于无法穷尽覆盖非预期输入,即便测试准确率达100%也不代表具备符号级推理能力。鉴于三段论是逻辑推理与人类理性基础,结果表明单纯增加数据与训练时间无法确保符号级逻辑推理。
原文摘要 · Abstract (English)
By promoting vectors to spheres and enabling explicit model construction, neural networks can perform symbolic-level syllogistic reasoning without training data. We identify two fundamental limitations that prevent conventional data-driven machine learning systems from achieving this capability: training data generated by the combination table cannot distinguish all 24 valid syllogism types, and end-to-end premise-to-conclusion mapping creates contradictory targets within neural components. Experiments with two representative conventional systems, GPT-5 using linguistic inputs and Euler Net using visual inputs, support this analysis. ChatGPT GPT-5 may reach 100% accuracy in syllogistic reasoning, but with hallucinations. Because the learning process terminates upon reaching 100% accuracy, the system cannot progress beyond empirical accuracy to symbolic level reasoning. Random test data reduced Euler Net's accuracy to 56%. Repeatedly expanding the training set increased its accuracy to 97%, with perfect performance on 8 syllogism types. However, because unintended inputs cannot be exhaustively covered, even 100% test accuracy does not imply symbolic-level reasoning. Since syllogistic reasoning underpins logical reasoning and human rationality, these results suggest that increasing data and training time alone cannot ensure symbolic level logical reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。