用自述数据训练大模型,实现无需特定任务数据的个体行为通用模拟。
LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- 基于访谈或问卷的自述数据构建生成式个体代理,无需针对每项任务单独训练。
- 联合使用访谈与问卷可达到86%的预测准确率,优于仅用人口统计信息的74%。
- 能跨领域预测性格、经济行为等,且减少不同种族/意识形态群体间的误差差距。
机器学习在有大量结构化数据且目标明确时能较好预测人类行为,但此类模型通常仅适用于特定结果,需为每个目标收集训练数据,限制了其在新领域的应用。我们测试大型语言模型(LLMs)是否可通过自述数据构建态度与行为模拟,即“生成式代理”,在无需任务特异性训练数据的情况下预测多种结果。基于来自1052名美国人的多样化全国样本数据,我们分别利用(i)两小时半结构化访谈(美国之声项目问卷),(ii)包含一般社会调查项目和大五人格量表的结构化问卷,或(iii)两者结合的数据构建代理。在保留的通用社会调查项目上,仅用访谈、仅用问卷及两者结合的代理分别达到参与者自身两周重测一致性基准的83%、82%和86%,而仅用人口统计信息的代理为74%。结合访谈与问卷获得最高准确率,但相较于单一来源提升有限,表明一旦模型在某领域获得足够证据,预测收益便趋于饱和。此外,这些代理还能预测人格特质、经济博弈行为与实验反应,且相较于仅用人口统计信息的代理,降低了不同种族与意识形态群体间的预测误差差异。结果表明,基于定性或定量自述数据的LLM代理可在无任务特异性训练数据的前提下,支持跨结果的个体通用模拟。
原文摘要 · Abstract (English)
Machine learning can predict human behavior well when substantial structured data are available for well-defined outcomes. Such models are typically outcome-specific, however, requiring training data for each target outcome, limiting their applicability to new domains. We test whether large language models (LLMs) can relax these requirements by using self-report data to build attitudinal and behavioral simulations, or "generative agents," that can predict responses across outcomes without outcome-specific training data. Using data from a diverse national sample of 1,052 Americans, we built agents from (i) two-hour, semi-structured interviews elicited using the American Voices Project interview schedule, (ii) structured surveys including General Social Survey items and the Big Five personality inventory, or (iii) both sources combined. On held-out General Social Survey items, interview-only, survey-only, and combined agents achieved accuracies equal to 83%, 82%, and 86% of participants' own two-week test-retest consistency benchmark, respectively, compared with 74% for demographics-only agents. Combining interviews and surveys produced the highest accuracy, though gains over either source alone were modest, suggesting that predictive benefits from data begin to asymptote once the model has observed sufficient evidence within a domain. We find that these agents also predict personality traits, economic-game behavior, and experimental responses, while reducing accuracy disparities across racial and ideological groups relative to demographics-only agents. Together, these results show that LLM agents grounded in qualitative or quantitative self-reports can support general-purpose simulation of individuals across outcomes, without requiring task-specific training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。