arXiv:2503.11614cs.CL2025-03ACL被引 4

用反事实数据训练大模型,减少推理中的偏见和幻觉。

Neutralizing Bias in LLM Reasoning using Entailment Graphs

  • 构建反事实推理数据,消除模型对命题记忆的依赖。
  • 在原始和去偏置数据集上,推理准确率均显著提升。
  • 适合研究大模型推理偏差与鲁棒性提升的研究者。

大语言模型常被认为具备自然语言推理(NLI)能力,但近期研究发现其在NLI任务中仍存在因确证偏见(attestation bias)导致的幻觉问题,即过度依赖命题记忆形成捷径。为此,我们设计了一种无监督框架,用于构建反事实推理数据并微调大模型以降低确证偏见。为衡量偏见缓解效果,我们构建了带有随机替换前提谓词但假设保持不变的对抗性偏见NLI数据集。大量实验表明,该框架能显著减少由确证偏见引发的幻觉。进一步评估在原始NLI数据集及其去偏置版本(原实体被随机替换)上的表现,结果一致显示,经本框架微调的模型在原始与去偏置数据集上均取得更优的推理性能。

原文摘要 · Abstract (English)

LLMs are often claimed to be capable of Natural Language Inference (NLI), which is widely regarded as a cornerstone of more complex forms of reasoning. However, recent works show that LLMs still suffer from hallucinations in NLI due to attestation bias, where LLMs overly rely on propositional memory to build shortcuts. To solve the issue, we design an unsupervised framework to construct counterfactual reasoning data and fine-tune LLMs to reduce attestation bias. To measure bias reduction, we build bias-adversarial variants of NLI datasets with randomly replaced predicates in premises while keeping hypotheses unchanged. Extensive evaluations show that our framework can significantly reduce hallucinations from attestation bias. Then, we further evaluate LLMs fine-tuned with our framework on original NLI datasets and their bias-neutralized versions, where original entities are replaced with randomly sampled ones. Extensive results show that our framework consistently improves inferential performance on both original and bias-neutralized NLI datasets.

大模型推理偏见消除自然语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。