用统计方法融合大模型生成数据与真实数据,提升市场调研准确性。
Large Language Models for Market Research: A Data-augmentation Approach
- 提出新数据增强框架,将大模型生成数据与真实数据结合建模。
- 实证显示可降低24.9%至79.8%的数据成本,误差更小。
- 适合需要降低成本又保持精度的市场研究团队使用。
大语言模型(LLMs)在自然语言处理中表现卓越,其生成类人文本的能力为市场研究带来新可能,尤其在联合分析中,可缓解传统问卷调查成本高、规模受限的问题。然而,现有研究指出,直接用大模型生成数据替代人类数据会引入显著偏差。本文提出一种新的统计数据增强方法,将大模型生成数据与真实数据有机结合,得到具有稳定性和渐近正态性的估计量。相比简单替换的朴素方法,该方法能有效缓解偏差。我们还给出了有限样本下的估计误差上界。通过新冠疫苗偏好和跑车选择两个实证研究验证:本方法可减少24.9%至79.8%的数据成本,且估计误差更小;而朴素方法因模型偏差无法节省数据。结果表明,大模型数据虽不能直接替代人类数据,但在稳健统计框架下可作为重要补充。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have transformed artificial intelligence by excelling in complex natural language processing tasks. Their ability to generate human-like text has opened new possibilities for market research, particularly in conjoint analysis, where understanding consumer preferences is essential but often resource-intensive. Traditional survey-based methods face limitations in scalability and cost, making LLM-generated data a promising alternative. However, while LLMs have the potential to simulate real consumer behavior, recent studies highlight a significant gap between LLM-generated and human data, with biases introduced when substituting between the two. In this paper, we address this gap by proposing a novel statistical data augmentation approach that efficiently integrates LLM-generated data with real data in conjoint analysis. This results in statistically robust estimators with consistent and asymptotically normal properties, in contrast to naive approaches that simply substitute human data with LLM-generated data, which can exacerbate bias. We further present a finite-sample performance bound on the estimation error. We validate our framework through an empirical study on COVID-19 vaccine preferences, demonstrating its superior ability to reduce estimation error and save data and costs by 24.9% to 79.8%. In contrast, naive approaches fail to save data due to the inherent biases in LLM-generated data compared to human data. Another empirical study on sports car choices validates the robustness of our results. Our findings suggest that while LLM-generated data is not a direct substitute for human responses, it can serve as a valuable complement when used within a robust statistical framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。