arXiv:2512.12552cs.AI2025-12被引 3

大模型在决策中会放大人类认知偏差,尤其越复杂的模型越易因过度思考出错。

Large Language Newsvendor: Decision Biases and Cognitive Mechanisms

  • 用动态新闻商问题测试大模型决策行为,发现其复制并放大了经典偏差。
  • 越复杂的GPT-4反而最不理性,而优化效率的GPT-4o表现接近最优。
  • 适合关注AI决策风险与人机协同的管理者阅读。

尽管大型语言模型(LLMs)日益融入商业决策,但其可能复制甚至放大人类认知偏差,带来显著但未被充分理解的风险,尤其在供应链管理等高风险运营场景中。本文通过动态多轮实验,使用GPT-4、GPT-4o和LLaMA-8B测试五种已知决策偏差。结果发现,大模型持续表现出经典的“订货过低/过高”偏差,并显著放大需求追逐行为,优于人类基准。分析揭示“智能悖论”:更复杂的GPT-4因过度思考导致最大非理性,而效率优化的GPT-4o表现近乎最优。由于这些偏差在提供最优公式后仍存在,说明其源于架构限制而非知识不足。管理启示:应根据任务选择模型;需加强人机协同监督以避免重大失误;设计结构化规则提示可有效抑制模型启发式倾向,提升决策可靠性。

原文摘要 · Abstract (English)

Problem definition: Although large language models (LLMs) are increasingly integrated into business decision making, their potential to replicate and even amplify human cognitive biases cautions a significant, yet not well-understood, risk. This is particularly critical in high-stakes operational contexts like supply chain management. To address this, we investigate the decision-making patterns of leading LLMs using the canonical newsvendor problem in a dynamic setting, aiming to identify the nature and origins of their cognitive biases. Methodology/results: Through dynamic, multi-round experiments with GPT-4, GPT-4o, and LLaMA-8B, we tested for five established decision biases. We found that LLMs consistently replicated the classic ``Too Low/Too High'' ordering bias and significantly amplified other tendencies like demand-chasing behavior compared to human benchmarks. Our analysis uncovered a ``paradox of intelligence'': the more sophisticated GPT-4 demonstrated the greatest irrationality through overthinking, while the efficiency-optimized GPT-4o performed near-optimally. Because these biases persist even when optimal formulas are provided, we conclude they stem from architectural constraints rather than knowledge gaps. Managerial implications: First, managers should select models based on the specific task, as our results show that efficiency-optimized models can outperform more complex ones on certain optimization problems. Second, the significant amplification of bias by LLMs highlights the urgent need for robust human-in-the-loop oversight in high-stakes decisions to prevent costly errors. Third, our findings suggest that designing structured, rule-based prompts is a practical and effective strategy for managers to constrain models' heuristic tendencies and improve the reliability of AI-assisted decisions.

大模型决策认知偏差供应链管理人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。