arXiv:2608.30413cs.AI2026-08中稿 · EMNLP

构建可生成信念更新对话的测试框架,研究大模型如何应对新证据

DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark

论文配图:DERELAB: Probing Defeasible Reasoning and Confirmation Bias in LLMs with a Generative Benchmark
图 1 · 摘自论文原文
  • 基于参数化图结构生成多轮信念更新对话,每步有形式验证真值
  • 9个模型中近半数在新证据下仍固守旧结论,体现确认偏误
  • 适合研究模型推理鲁棒性与认知偏差的学者使用

缺陷性推理是一种基于当前合理证据做出推断,但可在获得新证据时撤销的推理方式。尽管已有研究探讨语言模型在缺陷性推理中的表现,但现有数据集静态且覆盖非单调推理类别不足。我们提出DeReLab,一种生成框架,从参数化图结构生成涵盖默认与继承推理的多轮信念更新对话,每轮均有形式验证的真值,实现对模型面对支持性与反驳性证据时响应行为的受控测量。该生成机制为分离特定推理需求的实验设计提供测试平台。应用于确认偏误研究,我们评估了九个开源及专有大语言模型,发现几乎所有模型均系统性接受一致证据而抵制不一致更新,部分模型虽识别出证据削弱但仍未能修正结论。我们认为本工作及发现将推动未来对语言模型缺陷性推理能力的评估研究。

原文摘要 · Abstract (English)

Defeasible reasoning is a type of reasoning where inferences are drawn from plausible current evidence, but can be retracted upon the introduction of newer evidence. Although recent studies have examined language-model behaviors in defeasible reasoning, the datasets have been static and lack wide coverage of non-monotonic reasoning categories. We introduce DeReLab, a generative framework that produces multi-turn belief-updating conversations from parameterized graph structures across default and inheritance reasoning, with formally verified ground truth at every turn, enabling controlled measurement of how models respond to confirming and disconfirming evidence. This controlled generation process creates a testbed for experimental designs that isolate specific reasoning demands. Applying this capability to the study of confirmation bias, we evaluate nine open and proprietary large language models and find that nearly all exhibit a systematic tendency to accept congruent evidence while resisting incongruent updates, with several models correctly identifying a weakening update yet failing to revise their conclusion. We believe our work and findings will facilitate future research on evaluating language models in defeasible reasoning.

推理能力确认偏误大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。