arXiv:2604.01366cs.AI2026-04

发现大模型认知偏差可定位并干预,且不同偏差类型需不同应对策略。

CogBias: Measuring and Mitigating Cognitive Bias in Large Language Models

  • 构建四类认知偏差评测基准,识别模型内部编码的偏差方向。
  • 激活调控使偏差分数降低26%~32%,关键任务性能基本不受影响。
  • 跨模型偏差结构近正交但干预效果相似,暗示功能组织具有共性。

大型语言模型在高风险决策场景中应用日益广泛。尽管已有研究揭示模型存在行为层面的认知偏差,但这些偏差是否对应可识别的内部表征,以及能否通过针对性干预缓解,仍属未知。本文将大模型认知偏差定义为在具备可计算真值基准的任务中系统性、可复现的答案偏离,并提出LLM CogBias基准,涵盖判断、信息处理、社会和回应四类偏差。评估三个大模型发现,四类偏差均系统性出现,且其幅度与去偏响应显著依赖于偏差类型:提示层去偏可大幅减少回应偏差,却加剧判断偏差。通过对比设计下的线性探测,证实偏差以线性可分方向存在于模型激活空间。进一步应用激活调控,实现偏差分数26%~32%的下降,同时在25个下游任务上保持性能(Llama:几乎无退化;Qwen:判断偏差最多-19.0个百分点)。尽管不同模型间偏差表征近乎正交(平均余弦相似度0.01),但调控对各类模型的去偏速率相近(r(246)=0.621, p<0.001),表明其可能存在共享的功能组织。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly deployed in high-stakes decision-making contexts. While prior work has shown that LLMs exhibit cognitive biases behaviorally, whether these biases correspond to identifiable internal representations and can be mitigated through targeted intervention remains an open question. We define LLM cognitive bias as systematic, reproducible deviations from correct answers in tasks with computable ground-truth baselines, and introduce LLM CogBias, a benchmark organized around four families of cognitive biases: Judgment, Information Processing, Social, and Response. We evaluate three LLMs and find that cognitive biases emerge systematically across all four families, with magnitudes and debiasing responses that are strongly family-dependent: prompt-level debiasing substantially reduces Response biases but backfires for Judgment biases. Using linear probes under a contrastive design, we show that these biases are encoded as linearly separable directions in model activation space. Finally, we apply activation steering to modulate biased behavior, achieving 26--32\% reduction in bias score (fraction of biased responses) while preserving downstream capability on 25 benchmarks (Llama: negligible degradation; Qwen: up to $-$19.0pp for Judgment biases). Despite near-orthogonal bias representations across models (mean cosine similarity 0.01), steering reduces bias at similar rates across architectures ($r(246)$=.621, $p$<.001), suggesting shared functional organization.

大模型认知偏差去偏激活调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。