arXiv:2508.16910cs.CL2025-08中稿 · the 34th ACM Inter…被引 12

用因果干预方法消除大模型推理偏见,提升知识密集型任务准确率

Unbiased Reasoning for Knowledge-Intensive Tasks in Large Language Models via Conditional Front-Door Adjustment

  • 基于条件前门调整构建因果提示框架,模拟不同外部知识下的推理
  • 在多个大模型和基准数据集上,准确率显著优于现有方法
  • 适合需要高可靠性推理的场景,如医疗、法律等专业领域

大型语言模型在自然语言处理中表现出色,但在需要深度推理和外部知识整合的知识密集型任务上仍表现不佳。尽管检索增强生成(RAG)和思维链(CoT)等方法提升了模型对外部知识的利用能力,但其内部偏见仍常导致错误答案。本文提出一种新型因果提示框架——条件前门提示(CFD-Prompting),在给定外部知识条件下,实现查询与答案之间因果效应的无偏估计,有效缓解内部偏见。通过构建反事实外部知识,该框架模拟查询在不同上下文下的行为,解决查询固定无法直接干预的问题。相较于标准前门调整,条件变体假设更弱,增强了推理过程的鲁棒性和泛化能力。在多个大模型和基准数据集上的大量实验表明,CFD-Prompting 在准确率和鲁棒性上均显著优于现有基线方法。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown impressive capabilities in natural language processing but still struggle to perform well on knowledge-intensive tasks that require deep reasoning and the integration of external knowledge. Although methods such as Retrieval-Augmented Generation (RAG) and Chain-of-Thought (CoT) have been proposed to enhance LLMs with external knowledge, they still suffer from internal bias in LLMs, which often leads to incorrect answers. In this paper, we propose a novel causal prompting framework, Conditional Front-Door Prompting (CFD-Prompting), which enables the unbiased estimation of the causal effect between the query and the answer, conditional on external knowledge, while mitigating internal bias. By constructing counterfactual external knowledge, our framework simulates how the query behaves under varying contexts, addressing the challenge that the query is fixed and is not amenable to direct causal intervention. Compared to the standard front-door adjustment, the conditional variant operates under weaker assumptions, enhancing both robustness and generalisability of the reasoning process. Extensive experiments across multiple LLMs and benchmark datasets demonstrate that CFD-Prompting significantly outperforms existing baselines in both accuracy and robustness.

因果推理大模型知识融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。