arXiv:2606.21820cs.SIcs.AI2026-06综述

用大模型生成假问卷数据,辅助疫情决策研究。

Generating Public Health Responses using Survey-Augmented Large Language Models

论文配图:Generating Public Health Responses using Survey-Augmented Large Language Models
图 1 · 摘自论文原文
  • 用聚类分析识别疫苗态度群体,再以提示工程生成合成问卷。
  • 合成数据能还原人口统计特征和行为分布,但关联性较弱。
  • 适合用于建模初期探索,不替代真实调查。

流行病学模型常依赖调查数据来反映个体在疫苗接种或防护行为上的决策。然而,大规模重复调查成本高、耗时长,且难以覆盖多样情景。本文研究大语言模型(LLMs)能否生成再现真实人群模式的合成调查响应。基于流感路径(FluPaths)的纵向数据,我们通过聚类分析识别出对疫苗持积极或消极态度的群体,随后采用群体引导提示法评估多个LLM在多轮疫情中生成合成响应的效果。结果显示,合成数据总体上复现了真实数据中的年龄、性别等人口统计特征,以及疫苗信念、风险感知和健康行为的分布;但在个体内部这些因素的协同变化方面表现较差。部分模型更可靠地再现群体级疫苗接种趋势,但不同疫情波次间表现不一。我们还训练分类器区分真实与合成记录,发现生成数据仍可被识别为合成。总体表明,LLM生成的调查数据可能作为探索性数据增强工具,助力基于代理的流行病建模,但需进一步改进与验证后方可使用,不应直接替代人类调查数据。

原文摘要 · Abstract (English)

Epidemiological models often rely on survey data to represent how individuals make health-related decisions, such as whether to vaccinate or adopt protective behaviors. However, repeated large-scale surveys are costly, time-consuming, and limited in the range of scenarios they can capture. In this work, we investigate whether large language models (LLMs) can generate synthetic survey responses that reproduce patterns observed in real populations. Using longitudinal data from the FluPaths surveys, we first identify groups associated with broadly positive or negative attitudes toward vaccination through clustering analysis. We then evaluate several LLMs using a cluster-informed prompting approach to generate synthetic survey responses across multiple epidemic waves. Across models, the synthetic data generally reproduce the distributions of demographic characteristics, vaccination-related beliefs, risk perceptions, and health behaviors observed in the survey data. However, they are less successful at capturing how these factors vary together within respondents. Some models reproduce group-level vaccination trends more reliably than others, although performance varies across waves. We also trained a classifier to distinguish real from synthetic records and found that the generated responses remained identifiable as synthetic. Overall, our findings suggest that LLM-generated survey data may provide a useful tool for exploratory data augmentation and we hope that it could support agent-based epidemic modeling approaches. However, the generated data should not be treated as a substitute for human survey data without further methodological improvements and validation.

大模型生成公共健康问卷模拟流行病建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。