测试大模型在可撤销推理中的表现,揭示其逻辑能力边界。
Benchmarking Defeasible Reasoning with Large Language Models -- Initial Experiments and Future Directions
- 将可撤销规则转化为自然语言,构建适配LLM的推理基准
- 对比ChatGPT与形式化逻辑在非单调推理中的一致性,发现偏差
- 为评估模型推理可靠性提供新方法,适合关注AI可信性的研究者
大型语言模型(LLMs)因卓越性能在人工智能领域备受关注。因此,深入理解其能力与局限至关重要,尤其在非单调推理方面。本文提出一个对应多种可撤销规则推理模式的基准测试。我们通过将现有的可撤销逻辑推理基准中的规则转换为适合大语言模型处理的自然语言文本,开展了初步实验,利用ChatGPT评估其在非单调规则推理任务中的表现,并与基于可撤销逻辑定义的推理模式进行比较。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have gained prominence in the AI landscape due to their exceptional performance. Thus, it is essential to gain a better understanding of their capabilities and limitations, among others in terms of nonmonotonic reasoning. This paper proposes a benchmark that corresponds to various defeasible rule-based reasoning patterns. We modified an existing benchmark for defeasible logic reasoners by translating defeasible rules into text suitable for LLMs. We conducted preliminary experiments on nonmonotonic rule-based reasoning using ChatGPT and compared it with reasoning patterns defined by defeasible logic.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。