arXiv:2608.14320cs.AIcs.CL2026-08被引 1

测试大模型在不同锚点路径下的认知偏移,发现真实锚点影响更大且高精度模型仍易受骗。

AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs

论文配图:AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs
图 1 · 摘自论文原文
  • 设计多路径锚点评估框架,区分无关与合理锚点的影响力
  • 合理锚点在强路径下使判断偏离平均达10.3分以上,远离答案时影响减弱
  • 即使顶级模型准确率超95%,仍易被合理锚点误导

锚定效应是人类判断中因初始参考值而产生偏差的心理现象。近期研究表明大语言模型(LLMs)也存在类似行为。然而现有研究通常仅评估有限锚点路径,且很少区分无关与合理锚点。本文提出AnchorBench,一个基于显式锚点相关性维度的多路径锚定效应评估基准。在14个模型(包括10个开源模型和4个前沿API模型)及大量控制提示下,我们发现:(1) 锚定效应强烈依赖路径;(2) 当通过强路径引入时,合理锚点比无关锚点引发更大偏差;(3) 随着锚点与证据支持答案距离增加,影响普遍减弱,尤其在External和RAG任务中表现明显;(4) 即使在无锚条件下的高任务准确率(Acc$_{10}$:答案在黄金标准10分内)也不代表鲁棒性:即便前沿API模型控制准确率超过95%,仍对合理锚点敏感。

原文摘要 · Abstract (English)

The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther from the evidence-supported answer, most clearly on External and RAG, and (4) high task accuracy on the anchor-free control condition (Acc$_{10}$: answers within 10 points of gold) does not guarantee robustness: even frontier API models above 95% control accuracy remain susceptible to plausible anchors.

大模型偏见认知偏差评估基准锚定效应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。