为生成式代理模型提供机制合理性评估框架
Mechanism Plausibility in Generative Agent-Based Modeling

- 提出四层级机制可实现性评估标准
- 区分生成充分性与机制合理性,避免混淆
- 适合关注模型解释力的社科仿真研究者
大语言模型(LLMs)能在未显式编程规则的情况下生成多样化的高层现象,已被广泛应用于各类基于代理的模型(ABMs)和社会模拟中。尽管现有研究关注其生成特定现象的能力(如社交媒体人类行为或博弈论场景中的异类行为),但能力、预测与解释三者本质不同。根据科学哲学与机制理论,解释需揭示现象如何由相关实体及其活动组织产生。模型者若缺乏跨领域知识支撑,难以判断实验特征或模拟是否真正推动了解释进展。本文整合最新LLM-ABM研究与当代科学哲学成果,构建四层级机制可实现性评估体系,明确区分模型的生成充分性(能否复现现象)与机制可实现性(现象如何被生成),厘清预测型与解释型模型的不同角色,提出‘机制可实现性量表’。
原文摘要 · Abstract (English)
Large language models (LLMs) can generate high-level diverse phenomena without explicitly programmed rules. This capability has led to their adoption within different agent-based models (ABMs) and social simulations. Recent studies investigate their ability to generate different phenomena of interest, for example, human behavior on social media platforms or alien behavior in game-theoretic scenarios. However, capability, prediction, and explanation are different--drawing from the philosophy of science and mechanisms literature, explanation requires showing, to some degree, how a phenomenon is produced by related organized entities and activities. For modelers, describing the characteristics of an experiment or whether a simulation provides progress in capability (or explanation), can be difficult without being grounded in potentially distant research areas. We integrate recent work on LLM-ABMs with contemporary philosophy of science literature and use it to operationalize a definition of 'plausibility' in a four-level scale. Our scale separates the evaluation of a model's generative sufficiency (ability to reproduce a phenomenon) from its mechanistic plausibility (how the phenomenon could be produced), and clarifies the distinct roles of different models, such as predictive and explanatory ones. We introduce this as the Mechanism Plausibility Scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。