用大模型模拟物流行为,发现表面像人但决策逻辑不同。
Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research
- 用大模型生成人类行为响应,模拟外卖场景中的互动
- 6个大模型在957人实验中表现接近人类,但决策路径不一致
- 提出双层验证框架,适合研究者与从业者参考
基于大语言模型(LLM)的生成式智能体模型(GABMs)为物流与供应链管理(LSCM)实证研究提供了新可能,可模拟复杂人类行为。本文通过受控实验,在外卖配送场景中对比6个前沿大模型与957名参与者(477组配对)的表现,采用调节中介设计评估其有效性。结果表明,尽管部分大模型在行为表现上与人类具有表面等价性(通过两样本等效检验),但结构方程模型揭示部分模型存在人工决策路径,与真实人类行为不一致。研究提出双重验证框架:一是人类等价性测试,二是决策过程验证,强调需同时关注行为相似性与内在逻辑合理性。该框架为LSCM研究者提供严谨的GABM开发指南,并为从业者选择适用的大模型提供实证依据。
原文摘要 · Abstract (English)
Generative Agent-Based Models (GABMs) powered by large language models (LLMs) offer promising potential for empirical logistics and supply chain management (LSCM) research by enabling realistic simulation of complex human behaviors. Unlike traditional agent-based models, GABMs generate human-like responses through natural language reasoning, which creates potential for new perspectives on emergent LSCM phenomena. However, the validity of LLMs as proxies for human behavior in LSCM simulations is unknown. This study evaluates LLM equivalence of human behavior through a controlled experiment examining dyadic customer-worker engagements in food delivery scenarios. I test six state-of-the-art LLMs against 957 human participants (477 dyads) using a moderated mediation design. This study reveals a need to validate GABMs on two levels: (1) human equivalence testing, and (2) decision process validation. Results reveal GABMs can effectively simulate human behaviors in LSCM; however, an equivalence-versus-process paradox emerges. While a series of Two One-Sided Tests (TOST) for equivalence reveals some LLMs demonstrate surface-level equivalence to humans, structural equation modeling (SEM) reveals artificial decision processes not present in human participants for some LLMs. These findings show GABMs as a potentially viable methodological instrument in LSCM with proper validation checks. The dual-validation framework also provides LSCM researchers with a guide to rigorous GABM development. For practitioners, this study offers evidence-based assessment for LLM selection for operational tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。