arXiv:2501.08579cs.CL2025-01被引 28

LLM模拟人类行为不可靠,需改进数据与设计

LLM-based Human Simulations Have Not Yet Been Reliable

  • 分析现有模拟框架,指出模型与设计缺陷是问题根源
  • 提出增强数据、提升模型、优化设计的系统性解决方案
  • 提供可操作算法和资源库,助力可信模拟研究

大型语言模型(LLMs)在社会、经济、政策和心理等领域的行为模拟中应用日益广泛。然而,当前基于LLM的人类行为模拟仍缺乏可靠性,其结果与真实人类行为存在显著差异。本文系统回顾了上述领域中基于LLM的人类模拟方法,识别出其通用框架、最新进展及持续存在的局限性。研究发现,这些差异主要源于LLM固有缺陷与模拟设计中的问题。基于此,我们提出了一个系统性解决方案框架,强调丰富数据基础、提升模型能力与确保模拟设计稳健性。最后,我们设计了一套结构化算法以实现该框架,旨在指导可信且对齐人类行为的模拟实践。为促进后续研究,我们在https://github.com/Persdre/awesome-llm-human-simulation 提供了精选文献与资源列表。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.

人类模拟LLM可靠性系统框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。