研究提示词中的认知偏差如何影响大模型输出,揭示其误导性风险
Investigating the Effects of Cognitive Biases in Prompts on Large Language Model Outputs
- 在提示词中引入确认偏误等认知偏差,观察对模型输出的影响
- 微小偏差即显著改变大模型答案选择,尤其在金融问答中更明显
- 适合关注AI可靠性与提示工程优化的研究者和开发者
本文研究认知偏差对大语言模型(LLM)输出的影响。诸如确认偏误和可得性偏误等认知偏差会通过提示词扭曲用户输入,可能导致模型生成不忠实且具有误导性的输出。基于系统性框架,本研究将多种认知偏差引入提示词,并在多个基准数据集上评估其对模型准确率的影响,涵盖通用及金融问答场景。结果表明,即使微小的偏差也能显著改变大模型的答案选择,凸显了设计具备偏差意识的提示词与制定缓解策略的紧迫性。此外,注意力权重分析显示,这些偏差会改变模型内部决策过程,导致注意力分布异常,进而关联输出错误。该研究对AI开发者与使用者在提升各类应用场景下AI系统的鲁棒性与可靠性方面具有重要意义。
原文摘要 · Abstract (English)
This paper investigates the influence of cognitive biases on Large Language Models (LLMs) outputs. Cognitive biases, such as confirmation and availability biases, can distort user inputs through prompts, potentially leading to unfaithful and misleading outputs from LLMs. Using a systematic framework, our study introduces various cognitive biases into prompts and assesses their impact on LLM accuracy across multiple benchmark datasets, including general and financial Q&A scenarios. The results demonstrate that even subtle biases can significantly alter LLM answer choices, highlighting a critical need for bias-aware prompt design and mitigation strategy. Additionally, our attention weight analysis highlights how these biases can alter the internal decision-making processes of LLMs, affecting the attention distribution in ways that are associated with output inaccuracies. This research has implications for Al developers and users in enhancing the robustness and reliability of Al applications in diverse domains.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。