用大模型生成虚拟问卷,模拟德国人出行偏好。
Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?
- 基于人物画像的合成数据生成方法,融合人口与行为特征。
- 生成数据与真实调查(MiD 2017)在出行偏好上高度一致。
- 适合交通规划与社科研究,成本低且保护隐私。
本研究探讨大型语言模型(LLMs)生成人工问卷的潜力,聚焦德国个人出行偏好。通过利用LLM生成合成数据,旨在解决传统问卷方法存在的高成本、低效率及可扩展性差等问题。提出一种结合“人物画像”(即人口统计与行为特征组合)的新方法,并与五种其他合成调查方法进行对比,这些方法在使用真实数据和方法复杂度上各有差异。以德国全面出行调查数据集MiD 2017为基准,评估合成数据与真实模式的一致性。结果表明,LLM能有效捕捉人口属性与出行偏好之间的复杂依赖关系,同时具备探索假设情景的灵活性。该方法为交通规划与社会科学研究提供了可扩展、低成本且隐私保护的数据生成新路径。
原文摘要 · Abstract (English)
This study explores the potential of Large Language Models (LLMs) to generate artificial surveys, with a focus on personal mobility preferences in Germany. By leveraging LLMs for synthetic data creation, we aim to address the limitations of traditional survey methods, such as high costs, inefficiency and scalability challenges. A novel approach incorporating "Personas" - combinations of demographic and behavioural attributes - is introduced and compared to five other synthetic survey methods, which vary in their use of real-world data and methodological complexity. The MiD 2017 dataset, a comprehensive mobility survey in Germany, serves as a benchmark to assess the alignment of synthetic data with real-world patterns. The results demonstrate that LLMs can effectively capture complex dependencies between demographic attributes and preferences while offering flexibility to explore hypothetical scenarios. This approach presents valuable opportunities for transportation planning and social science research, enabling scalable, cost-efficient and privacy-preserving data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。