无需标注数据,用双重反事实检验提升大模型因果推理能力
Better Think Thrice: Learning to Reason Causally with Double Counterfactual Consistency
- 引入双重反事实一致性(DCC)方法,在推理时评估因果干预与反事实预测能力
- 在多个任务上验证主流大模型因果推理缺陷,且不依赖人工标注的反事实数据
- 可作为无训练测试时筛选策略,直接提升多类模型的推理表现
尽管大型语言模型在推理基准上表现强劲,但在面对反事实问题时仍显脆弱,暴露出因果推理能力不足。现有研究虽证明标注的反事实任务可作为评估工具,但大规模生成此类数据受限。本文提出一种轻量级推理时方法——双重反事实一致性(DCC),无需标注数据即可检验模型执行因果干预和反事实预测两项关键因果推理能力。通过DCC,我们评估了多种主流大模型在不同推理任务和干预下的因果推理能力,并证明其可作为无训练的测试时拒绝采样准则,有效提升多个模型家族在推理任务上的表现。
原文摘要 · Abstract (English)
Despite their strong performance on reasoning benchmarks, large language models (LLMs) have proven brittle when presented with counterfactual questions, suggesting weaknesses in their causal reasoning ability. While recent work has demonstrated that labeled counterfactual tasks can be useful benchmarks of LLMs' causal reasoning, producing such data at the scale required to cover the vast potential space of counterfactuals is limited. In this work, we introduce double counterfactual consistency (DCC), a lightweight inference-time method for measuring and guiding the ability of LLMs to reason causally. Without requiring labeled counterfactual data, DCC verifies a model's ability to execute two important elements of causal reasoning: causal intervention and counterfactual prediction. Using DCC, we evaluate the causal reasoning abilities of various leading LLMs across a range of reasoning tasks and interventions. Moreover, we demonstrate the effectiveness of DCC as a training-free test-time rejection sampling criterion and show that it can directly improve performance on reasoning tasks across multiple model families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。