arXiv:2409.00103cs.CLcs.AI2024-09AAAI被引 3

研究大模型在因果推理中对细微差异的判断一致性

Nuance Matters: Probing Epistemic Consistency in Causal Reasoning

  • 提出因果认知一致性概念,评估模型对细微因果中间项的判断是否自洽
  • 21个主流大模型测试显示,多数模型在极性与强度判断上不一致
  • 发现内部词元概率可辅助提升判断一致性,适合关注模型可靠性研究者

为填补这一空白,本研究引入因果认知一致性概念,聚焦大语言模型(LLMs)在区分因果推理中具有细微差异的中间项时的自我一致性。我们提出一套新度量指标——强度排序一致性、跨组位置一致性与组内聚类度,用于评估该能力。通过对21个高影响力大模型(包括GPT-4、Claude3和LLaMA3-70B)的广泛实证研究,我们获得有力证据表明,当前模型在识别因果推理中中间项的极性和强度时难以维持认知一致性。此外,我们探索了利用内部词元概率作为辅助工具以增强因果认知一致性的潜力。综上所述,本研究通过探究因果推理中精细中间项上的自我一致性,弥补了人工智能研究中的关键空白。

原文摘要 · Abstract (English)

To address this gap, our study introduces the concept of causal epistemic consistency, which focuses on the self-consistency of Large Language Models (LLMs) in differentiating intermediates with nuanced differences in causal reasoning. We propose a suite of novel metrics -- intensity ranking concordance, cross-group position agreement, and intra-group clustering -- to evaluate LLMs on this front. Through extensive empirical studies on 21 high-profile LLMs, including GPT-4, Claude3, and LLaMA3-70B, we have favoring evidence that current models struggle to maintain epistemic consistency in identifying the polarity and intensity of intermediates in causal reasoning. Additionally, we explore the potential of using internal token probabilities as an auxiliary tool to maintain causal epistemic consistency. In summary, our study bridges a critical gap in AI research by investigating the self-consistency over fine-grained intermediates involved in causal reasoning.

因果推理大模型评估认知一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。