arXiv:2506.07064cs.CLcs.AI2025-06ACL被引 3

构建复杂常识推理基准Com²,测试大模型对隐含因果关系的深层理解能力。

Com$^2$: A Causal-Guided Benchmark for Exploring Complex Commonsense Reasoning in Large Language Models

  • 用因果事件图结构化复杂常识,通过干预生成不同情景
  • 大模型在深度和广度推理上表现差,慢思考能部分缓解
  • 适合研究常识推理、因果建模与大模型可解释性的学者

大语言模型(LLMs)通过预训练掌握了大量简单显性常识,能在基础常识推理中达到人类水平。然而,在依赖简单常识推导出的复杂隐性常识(如事件长期影响)方面,大模型仍表现不佳,这正是人类更关注的层面。现有研究多聚焦数学与代码任务,复杂常识推理因不确定性和缺乏结构而被忽视。为此,我们提出基准Com²,聚焦复杂常识推理。首先利用因果事件图构建结构化复杂常识;再基于因果理论(如干预)修改因果图,生成符合人类关切的不同情景;最后用大模型以慢思考方式生成示例,受修改后因果图逻辑关系引导。此外,我们还引入侦探故事构建更具挑战性的子集。实验表明,大模型在推理深度与广度上表现不足,但后训练和慢思考可有效改善。代码与数据已公开于https://github.com/Waste-Wood/Com2。

原文摘要 · Abstract (English)

Large language models (LLMs) have mastered abundant simple and explicit commonsense knowledge through pre-training, enabling them to achieve human-like performance in simple commonsense reasoning. Nevertheless, LLMs struggle to reason with complex and implicit commonsense knowledge that is derived from simple ones (such as understanding the long-term effects of certain events), an aspect humans tend to focus on more. Existing works focus on complex tasks like math and code, while complex commonsense reasoning remains underexplored due to its uncertainty and lack of structure. To fill this gap and align with real-world concerns, we propose a benchmark Com$^2$ focusing on complex commonsense reasoning. We first incorporate causal event graphs to serve as structured complex commonsense. Then we adopt causal theory~(e.g., intervention) to modify the causal event graphs and obtain different scenarios that meet human concerns. Finally, an LLM is employed to synthesize examples with slow thinking, which is guided by the logical relationships in the modified causal graphs. Furthermore, we use detective stories to construct a more challenging subset. Experiments show that LLMs struggle in reasoning depth and breadth, while post-training and slow thinking can alleviate this. The code and data are available at https://github.com/Waste-Wood/Com2.

常识推理因果建模大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。