arXiv:2602.22070cs.AI2026-02

大模型对人类与算法的偏见不一致,表面信任人,实际更选算法。

Language Models Exhibit Inconsistent Biases Towards Algorithmic Agents and Human Experts

  • 用人类偏好和真实表现两种方式测试模型决策倾向。
  • 看到算法表现差时,仍更倾向选择算法下注。
  • 揭示模型在高风险场景中存在不可靠的偏见行为。

大型语言模型越来越多地用于需整合人类专家与算法代理信息的决策任务。本文借鉴行为经济学实验范式,评估8个不同LLM在将决策委托给人类专家或算法代理时的表现。研究采用两种任务呈现方式:一是直接询问对两类代理的信任度(陈述偏好),二是提供两者的实际表现示例并要求做出有激励的抉择(揭示偏好)。结果显示,当被问及信任度时,模型更信任人类专家,这与人类结果一致;但当面对实际表现对比并需下注时,即使算法表现更差,模型仍更倾向于选择算法。这一矛盾表明,模型对人类与算法存在不一致的偏见,需在高风险场景中谨慎对待。此外,模型对任务呈现方式敏感,提示应重视AI安全评估的稳健性。

原文摘要 · Abstract (English)

Large language models are increasingly used in decision-making tasks that require them to process information from a variety of sources, including both human experts and other algorithmic agents. How do LLMs weigh the information provided by these different sources? We consider the well-studied phenomenon of algorithm aversion, in which human decision-makers exhibit bias against predictions from algorithms. Drawing upon experimental paradigms from behavioural economics, we evaluate how eightdifferent LLMs delegate decision-making tasks when the delegatee is framed as a human expert or an algorithmic agent. To be inclusive of different evaluation formats, we conduct our study with two task presentations: stated preferences, modeled through direct queries about trust towards either agent, and revealed preferences, modeled through providing in-context examples of the performance of both agents. When prompted to rate the trustworthiness of human experts and algorithms across diverse tasks, LLMs give higher ratings to the human expert, which correlates with prior results from human respondents. However, when shown the performance of a human expert and an algorithm and asked to place an incentivized bet between the two, LLMs disproportionately choose the algorithm, even when it performs demonstrably worse. These discrepant results suggest that LLMs may encode inconsistent biases towards humans and algorithms, which need to be carefully considered when they are deployed in high-stakes scenarios. Furthermore, we discuss the sensitivity of LLMs to task presentation formats that should be broadly scrutinized in evaluation robustness for AI safety.

大模型偏见决策机制可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。