用大模型生成消费者洞察数据,省时省钱还可用。
Synthetic Consumer Insight Generation with Large Language Models
- 用大模型模拟消费者对城市旅游的反应,替代真实调研。
- 模型生成内容与真人回答在主题上高度重合,但语言风格不同。
- 适合营销研究快速获取数据,但需注意语言表达差异。
现代数据驱动营销依赖大量消费者数据,但收集成本高、耗时长且难以规模化。本研究探讨大型语言模型(LLMs)是否可用于生成项目法(projective techniques)所需的合成消费者数据,此类方法旨在揭示消费者的联想、情感、需求与愿望。我们在多个项目任务、大模型、提示策略及温度设置下测试了LLM生成的回应,并与一项关于城市旅游目的地感知的原始研究中的人类响应进行对比。通过语言学分析、多样性与集中度指标、主题模型和关键词分析,发现人类与LLM响应在主要话题和关联性上存在显著重叠,但在风格、语言结构以及多样性生成方式上存在重要差异。研究提出使用建议:如何最优利用LLM生成合成消费者数据,模型与提示选择如何影响响应质量,以及识别其局限性。
原文摘要 · Abstract (English)
Modern data-driven marketing relies on large amounts of consumer data, yet collecting such data can be costly, time-consuming, and difficult to scale. This research examines whether large language models (LLMs) can be used to generate synthetic consumer data for projective techniques, a set of methods designed to elicit consumer associations, emotions, wants, and needs. We test LLM-generated responses across multiple projective tasks, LLMs, prompting strategies, and temperature settings, and compare them with human responses from a primary research study on perceptions of city tourism destinations. Human and LLM responses were analyzed using linguistic measures, diversity and concentration metrics, topic models, and top-term analyses. The results show substantial overlap between human and LLM responses in broad topics and associations, but also important differences in style, linguistic structure, and the way diversity is generated. Recommendations are given on how to best utilize LLMs for generating synthetic consumer data, how model and prompt choices shape response quality, and on recognizing the limitations of LLM synthetic consumer data generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。