发现大模型在相关性判断中受阈值启动效应影响,像人一样产生认知偏差。
AI Can Be Cognitively Biased: An Exploratory Study on Threshold Priming in LLM-Based Batch Relevance Assessment
- 通过不同相关分、批次长度和模型测试,验证了阈值启动效应的存在
- 无论模型或参数规模,后期文档评分受前期影响,高分后易降分
- 提示研究者需关注大模型在信息检索中的类人认知偏差
认知偏差是系统性思维偏误,导致非理性判断与决策,已在多个领域被广泛研究。近期大型语言模型(LLMs)展现出先进理解能力,但可能从训练数据中继承人类偏见。尽管社会偏见已被充分研究,认知偏见仍关注不足,现有研究多集中于特定场景。我们探究了大模型在相关性判断中是否受阈值启动效应影响,这是信息检索(IR)领域的核心任务之一。启动效应指先前刺激无意识地影响后续行为与决策。实验使用TREC 2019 Deep Learning passage track的10个主题,测试不同文档相关分、批次长度及模型(GPT-3.5、GPT-4、LLaMa2-13B、LLaMa2-70B)下的AI判断。结果表明,无论组合或模型如何,若前序文档相关性高,则后续文档评分倾向降低,反之亦然。研究证实大模型判断与人类一致,受阈值启动偏误影响,提示研究者与工程师在设计、评估和审计大模型时应考虑潜在的人类认知偏差。
原文摘要 · Abstract (English)
Cognitive biases are systematic deviations in thinking that lead to irrational judgments and problematic decision-making, extensively studied across various fields. Recently, large language models (LLMs) have shown advanced understanding capabilities but may inherit human biases from their training data. While social biases in LLMs have been well-studied, cognitive biases have received less attention, with existing research focusing on specific scenarios. The broader impact of cognitive biases on LLMs in various decision-making contexts remains underexplored. We investigated whether LLMs are influenced by the threshold priming effect in relevance judgments, a core task and widely-discussed research topic in the Information Retrieval (IR) coummunity. The priming effect occurs when exposure to certain stimuli unconsciously affects subsequent behavior and decisions. Our experiment employed 10 topics from the TREC 2019 Deep Learning passage track collection, and tested AI judgments under different document relevance scores, batch lengths, and LLM models, including GPT-3.5, GPT-4, LLaMa2-13B and LLaMa2-70B. Results showed that LLMs tend to give lower scores to later documents if earlier ones have high relevance, and vice versa, regardless of the combination and model used. Our finding demonstrates that LLM%u2019s judgments, similar to human judgments, are also influenced by threshold priming biases, and suggests that researchers and system engineers should take into account potential human-like cognitive biases in designing, evaluating, and auditing LLMs in IR tasks and beyond.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。