发现大模型普遍存在锚定效应,且传统方法难消除。
Understanding the Anchoring Effect of LLM with Synthetic Data: Existence, Mechanism, and Potential Mitigations
- 用合成数据集测试大模型锚定效应存在性
- 浅层激活导致锚定偏差,常规策略无效
- 推理过程可部分缓解,适合偏见研究者参考
大型语言模型(LLMs)如ChatGPT的兴起推动了自然语言处理的发展,但认知偏见问题日益突出。本文研究锚定效应——一种认知偏差,即大脑过度依赖初始信息作为判断基准。我们探究了LLMs是否受锚定影响、其内在机制及潜在缓解策略。为支持大规模研究,我们引入新数据集SynAnchors(https://huggingface.co/datasets/TimTargaryen/SynAnchors)。结合优化评估指标,对当前广泛使用的多个LLMs进行基准测试。结果表明,LLMs普遍表现出锚定偏差,且主要由浅层网络激活引起;常规缓解策略无法有效消除该偏差,而推理能力可提供一定缓解作用。
原文摘要 · Abstract (English)
The rise of Large Language Models (LLMs) like ChatGPT has advanced natural language processing, yet concerns about cognitive biases are growing. In this paper, we investigate the anchoring effect, a cognitive bias where the mind relies heavily on the first information as anchors to make affected judgments. We explore whether LLMs are affected by anchoring, the underlying mechanisms, and potential mitigation strategies. To facilitate studies at scale on the anchoring effect, we introduce a new dataset, SynAnchors (https://huggingface.co/datasets/TimTargaryen/SynAnchors). Combining refined evaluation metrics, we benchmark current widely used LLMs. Our findings show that LLMs' anchoring bias exists commonly with shallow-layer acting and can not be eliminated by conventional strategies, while reasoning can offer some mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。