arXiv:2510.26606cs.AIcs.CL2025-10中稿 · EMNLP被引 3

对比大模型在规范与认知模态推理中的表现,发现其规范推理存在逻辑不一致和认知偏差。

Normative Reasoning in Large Language Models: A Comparative Benchmark from Logical and Modal Perspectives

  • 构建新数据集,对比规范与认知模态的推理模式。
  • 模型在规范推理中出现明显不一致性,类似人类认知偏差。
  • 适合关注AI伦理、逻辑推理可靠性的研究者参考。

规范推理涉及义务、允许等规范模态。尽管大语言模型在多种推理任务中表现优异,但其在规范推理方面的能力仍待深入探索。本文从逻辑与模态视角系统评估大模型在规范领域中的推理能力。为检验模型对规范模态的理解,我们将其与共享形式结构的认知模态推理进行对比。为此,我们构建了一个涵盖规范与认知领域广泛形式推理模式的新数据集,并引入影响人类推理的非形式认知因素。结果表明,尽管模型总体遵循有效推理模式,但在特定规范推理类型中表现出显著不一致,且呈现与心理学研究中观察到的人类认知偏差相似的行为。这些发现揭示了大模型在规范推理中实现逻辑一致性的挑战,并为提升其可靠性提供了洞见。所有数据与代码已公开于 https://github.com/kmineshima/NeuBAROCO。

原文摘要 · Abstract (English)

Normative reasoning is a type of reasoning that involves normative or deontic modality, such as obligation and permission. While large language models (LLMs) have demonstrated remarkable performance across various reasoning tasks, their ability to handle normative reasoning remains underexplored. In this paper, we systematically evaluate LLMs' reasoning capabilities in the normative domain from both logical and modal perspectives. Specifically, to assess how well LLMs reason with normative modals, we make a comparison between their reasoning with normative modals and their reasoning with epistemic modals, which share a common formal structure. To this end, we introduce a new dataset covering a wide range of formal patterns of reasoning in both normative and epistemic domains, while also incorporating non-formal cognitive factors that influence human reasoning. Our results indicate that, although LLMs generally adhere to valid reasoning patterns, they exhibit notable inconsistencies in specific types of normative reasoning and display cognitive biases similar to those observed in psychological studies of human reasoning. These findings highlight challenges in achieving logical consistency in LLMs' normative reasoning and provide insights for enhancing their reliability. All data and code are released publicly at https://github.com/kmineshima/NeuBAROCO.

规范推理大模型认知偏差逻辑一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。