发现大模型存在锚定偏差,且可量化分析其影响机制。
Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs
- 通过概率分布分析,揭示锚点改变模型输出整体分布。
- 用Shapley值精确计算锚点对模型决策的影响程度。
- 提出统一评分体系,适合评估模型安全与可解释性。
大型语言模型(LLMs)日益被视作行为主体与决策系统,但其表现出的认知偏差是否源于表层模仿或深层概率变化仍不明确。锚定偏差作为经典人类判断偏误,是关键检验案例。尽管已有研究显示LLMs存在锚定效应,但多数证据依赖表面输出,缺乏对内部机制与归因贡献的深入分析。本文通过三项贡献推进该研究:(1) 基于对数概率的行为分析,表明锚点会改变整个输出分布,且控制了训练数据污染;(2) 在结构化提示字段上使用精确的Shapley值归因,量化锚点对模型对数概率的影响;(3) 构建统一的锚定偏差敏感度评分,整合行为与归因证据,覆盖六种开源模型。结果显示,Gemma-2B、Phi-2和Llama-2-7B表现出显著锚定效应,归因信号显示锚点引发重加权。较小模型如GPT-2、Falcon-RW-1B和GPT-Neo-125M则表现不一,暗示规模可能调节敏感性。归因效果随提示设计而异,凸显将LLMs视为人类替代品的脆弱性。结果表明,LLMs中的锚定偏差具有鲁棒性、可测量性和可解释性,同时揭示应用中的潜在风险。更广泛而言,该框架连接了行为科学、大模型安全与可解释性,为评估其他认知偏差提供了可复现路径。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly examined as both behavioral subjects and decision systems, yet it remains unclear whether observed cognitive biases reflect surface imitation or deeper probability shifts. Anchoring bias, a classic human judgment bias, offers a critical test case. While prior work shows LLMs exhibit anchoring, most evidence relies on surface-level outputs, leaving internal mechanisms and attributional contributions unexplored. This paper advances the study of anchoring in LLMs through three contributions: (1) a log-probability-based behavioral analysis showing that anchors shift entire output distributions, with controls for training-data contamination; (2) exact Shapley-value attribution over structured prompt fields to quantify anchor influence on model log-probabilities; and (3) a unified Anchoring Bias Sensitivity Score integrating behavioral and attributional evidence across six open-source models. Results reveal robust anchoring effects in Gemma-2B, Phi-2, and Llama-2-7B, with attribution signaling that the anchors influence reweighting. Smaller models such as GPT-2, Falcon-RW-1B, and GPT-Neo-125M show variability, suggesting scale may modulate sensitivity. Attributional effects, however, vary across prompt designs, underscoring fragility in treating LLMs as human substitutes. The findings demonstrate that anchoring bias in LLMs is robust, measurable, and interpretable, while highlighting risks in applied domains. More broadly, the framework bridges behavioral science, LLM safety, and interpretability, offering a reproducible path for evaluating other cognitive biases in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。