用经典实验设计让大模型社会智能更可控,精准模拟认知偏差。
CoBRA: Programming Cognitive Bias in Social Agents Using Classic Social Science Experiments
- 将经典社会实验转为可复用的仿真环境,量化并控制智能体偏差。
- 通过闭环系统实现跨模型一致的行为表现,提升可重复性。
- 适合研究社会认知、人机交互与伦理对齐的学者使用。
本文提出CoBRA,一个用于在基于大语言模型的社会仿真中系统化设定代理行为的新工具包。传统方法通过自然语言隐式描述行为,常导致不同模型间行为不一致,且难以捕捉描述中的细微差别。CoBRA引入一种模型无关的控制机制,使研究者能显式指定所需行为特征,并确保跨模型的一致性。其核心是双组件闭环系统:(1) 认知偏差指数(Cognitive Bias Index),通过一组经验证的经典社会科学实验量化智能体表现出的认知偏差;(2) 行为调节引擎(Behavioral Regulation Engine),引导智能体行为以展现受控的认知偏差。通过CoBRA,我们展示了如何将已验证的社会科学知识(即经典实验)转化为可复用的“gym”仿真环境,该方法可推广至更丰富的社会与情感模拟场景。
原文摘要 · Abstract (English)
This paper introduces CoBRA, a novel toolkit for systematically specifying agent behavior in LLM-based social simulation. We found that conventional approaches that specify agent behavior through implicit natural-language descriptions often do not yield consistent behavior across models, and the resulting behavior does not capture the nuances of the descriptions. In contrast, CoBRA introduces a model-agnostic way to control agent behavior that lets researchers explicitly specify desired nuances and obtain consistent behavior across models. At the heart of CoBRA is a novel closed-loop system primitive with two components: (1) Cognitive Bias Index that measures the demonstrated cognitive bias of a social agent, by quantifying the agent's reactions in a set of validated classic social science experiments; (2) Behavioral Regulation Engine that aligns the agent's behavior to exhibit controlled cognitive bias. Through CoBRA, we show how to operationalize validated social science knowledge (i.e., classical experiments) as reusable "gym" environments for AI -- an approach that may generalize to richer social and affective simulations beyond bias alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。