用大模型生成虚拟人口问卷,模拟真实社会调研。
Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation

- 通过人物背景故事控制大模型生成,实现人口可调的问卷模拟。
- 在政治与医疗议题上,生成结果比传统方法更贴近真实调查数据。
- 开源易用,适合社科研究者快速测试问卷设计。
我们提出 Anamnesis,一个基于大语言模型的交互式系统,支持以人口统计特征可控的方式进行问卷模拟。该系统开源且面向非技术人员,可在虚拟人群上原型化并压力测试调查工具,而无需真实人类参与。平台整合了 Anthology 与 Alterity 框架,利用结构化叙事背景故事引导模型输出,在统一网页界面中支持开放式生成、概率性人口重采样及多模态(图像与音频)问卷。通过两个案例验证:(1) 复现皮尤研究中心美国趋势面板(ATP)在政治类型与生物医学议题上的部分结果;(2) 模拟《纽约客》标题竞赛中的人类偏好。结果显示,Anamnesis 生成的观点分布比标准角色提示基线更接近真实数据,提供透明、可复现且开源的替代方案,优于现有商业仿真服务。
原文摘要 · Abstract (English)
We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of survey instruments on virtual populations rather than real human subjects. The platform operationalizes the recently introduced Anthology and Alterity frameworks, which use structured narrative backstories to condition model responses, within a unified web interface. It supports open-ended generation, probabilistic demographic resampling, and multimodal (image and audio) surveys. We evaluate the system through two case studies: (1) replicating segments of Pew Research Center's American Trends Panel (ATP) on political typology and biomedical issues and (2) emulating human preference in the New Yorker Caption Contest. In both cases, Anamnesis produces opinion distributions that more closely match real-world survey data than standard persona-prompting baselines, offering a transparent, reproducible, and open-source alternative to proprietary simulation services.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。