研究自然语言可满足性问题分布,评估大模型推理能力边界
Natural Language Satisfiability: Exploring the Problem Distribution and Evaluating Transformer-based Language Models
- 分析不同复杂度类与语法结构对模型推理的影响
- 发现模型在高复杂度问题上准确率下降超过40%
- 为评测大模型逻辑推理提供更科学的问题分布基准
近年来,基于Transformer的语言模型在自然语言推理任务中取得显著进展。其中,判断命题可满足性是最基础且可归约至多数任务的核心问题。然而,从逻辑角度看,可满足性问题在多种维度上存在差异,可能影响模型的学习能力。自然语言中的可满足性实例因其表达语言片段的不同,可属于不同的计算复杂度类。尽管已有研究探讨该问题,但复杂度差异的影响尚未得到充分讨论。本文系统研究了不同计算复杂度类及语法构造对语言模型推理规则学习能力的影响,并通过实证研究探索了自然语言可满足性问题的分布特征,为真实评估语言模型提供了数据基础。
原文摘要 · Abstract (English)
Efforts to apply transformer-based language models (TLMs) to the problem of reasoning in natural language have enjoyed ever-increasing success in recent years. The most fundamental task in this area to which nearly all others can be reduced is that of determining satisfiability. However, from a logical point of view, satisfiability problems vary along various dimensions, which may affect TLMs' ability to learn how to solve them. The problem instances of satisfiability in natural language can belong to different computational complexity classes depending on the language fragment in which they are expressed. Although prior research has explored the problem of natural language satisfiability, the above-mentioned point has not been discussed adequately. Hence, we investigate how problem instances from varying computational complexity classes and having different grammatical constructs impact TLMs' ability to learn rules of inference. Furthermore, to faithfully evaluate TLMs, we conduct an empirical study to explore the distribution of satisfiability problems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。