arXiv:2608.05166cs.CLcs.CY2026-08

研究大模型在对话中如何被用户偏见影响,发现多数模型会放大偏见。

Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning

论文配图:Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning
图 1 · 摘自论文原文
  • 设计三条件实验框架,分离用户偏见与语义内容的影响
  • 8个前沿模型中6个在有偏对话中表现出更高偏见水平
  • 揭示偏见暴露与明确提示带来的两种相反行为机制

我们评估了先进指令微调大模型在真实多轮交互场景下的认知偏见表现。提出一种新的三条件实验框架,将用户偏见话语的影响与语义内容分开;构建包含24,300条经陪审团验证的用户提示的基准数据集,覆盖9×9目标-人类偏见交互矩阵全部81个单元。在8个前沿大模型中,发现有偏对话上下文相较于零样本基线,使6个模型的偏见表达系统性增强。我们识别出两种竞争性行为动态:接触偏见推理通常会放大后续偏见倾向,而显式偏见提示则常触发对齐相关抑制行为,降低表层偏见表达。我们公开框架、代码库和数据集,以支持未来对大模型上下文条件偏见与行为适应的研究。

原文摘要 · Abstract (English)

We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles the effect of exposure to a biased user turn from the effect of the turn's semantic content, alongside a benchmark of 24,300 jury-validated user prompts spanning all 81 cells of a 9x9 target-human bias interaction matrix. Across eight frontier LLMs, we find that biased conversational context systematically increases bias expression relative to zero-shot baselines in 6 of 8 models. We identify two competing behavioral dynamics underlying this effect: conversational exposure to biased reasoning generally amplifies downstream bias tendencies, while explicitly stated bias cues often trigger alignment-related suppression behaviors that reduce overt bias expression. We release our framework, codebase, and dataset to support future research on context-conditioned cognitive biases and behavioral adaptation in LLMs.

大模型偏见对话系统认知偏差行为机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。