arXiv:2510.13417cs.AIcs.CL2025-10被引 4

测试大模型在气候争论中发现隐含因果链的能力

Assessing LLM Reasoning Through Implicit Causal Chain Discovery in Climate Discourse

  • 让大模型补全气候议题中因果间的中间环节
  • 模型生成链条逻辑通顺但多靠模式匹配而非真推理
  • 适合研究因果推理与论辩分析的学者参考

大语言模型如何理解因果关系中的中间步骤?本研究通过隐含因果链发现任务,评估大模型在气候争议语境下的机制性因果推理能力。基于论辩研究中的极化气候讨论数据,我们要求九个大模型为给定的因果对生成所有可能的中间因果步骤。分析显示,模型生成的步骤数量和粒度存在差异;虽普遍自信且自洽,但其判断主要依赖关联模式匹配,而非真实因果推理。然而,人工评估确认了生成链条的逻辑连贯性与完整性。本研究提出的基准方法、诊断评估见解及包含因果链的基准数据集,为未来论辩场景中隐含机制性因果推理的研究奠定了坚实基础。

原文摘要 · Abstract (English)

How does a cause lead to an effect, and which intermediate causal steps explain their connection? This work scrutinizes the mechanistic causal reasoning capabilities of large language models (LLMs) to answer these questions through the task of implicit causal chain discovery. In a diagnostic evaluation framework, we instruct nine LLMs to generate all possible intermediate causal steps linking given cause-effect pairs in causal chain structures. These pairs are drawn from recent resources in argumentation studies featuring polarized discussion on climate change. Our analysis reveals that LLMs vary in the number and granularity of causal steps they produce. Although they are generally self-consistent and confident about the intermediate causal connections in the generated chains, their judgments are mainly driven by associative pattern matching rather than genuine causal reasoning. Nonetheless, human evaluations confirmed the logical coherence and integrity of the generated chains. Our baseline causal chain discovery approach, insights from our diagnostic evaluation, and benchmark dataset with causal chains lay a solid foundation for advancing future work in implicit, mechanistic causal reasoning in argumentation settings.

因果推理大模型评估论辩分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。