用AI生成数据提升市场研究精度,降低人工标注成本
Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence

- 将LLM输出作为特征融合进人类标签模型,避免直接替代
- 在联合分析中误差减半,标签需求降低75%以上
- 适合需降本增效的市场研究、健康保险等场景
市场研究常依赖昂贵的人类生成数据(如联合分析问卷、购买决策、实地实验结果)。近期大语言模型(LLMs)等AI系统可提供低成本辅助数据,但其输出并非目标结果的直接观测值,而是高维表示,与人类标签间关系复杂且未知。传统方法将AI预测视为真实标签的直接代理,当二者关系弱或设定错误时效率低下甚至不可靠。本文提出生成增强推断(GAI)框架,将AI生成内容作为信息性特征用于估计人类标签模型。GAI采用正交矩构造,支持在非参数、灵活的关系下实现一致估计和有效推断。我们建立了渐近正态性及关键主导性结果:在随机标注下,GAI在统一的无偏估计器类中最优,且在温和信息性条件下严格优于现有方法;即使标注样本不具代表性,扩展版GAI仍优于加权仅人类数据估计器。实证表明,GAI在多种市场研究场景中超越基线。联合分析中,估计误差减半,人类标注需求减少超75%;定价研究中,所有方法使用相同辅助输入时,GAI持续领先;健康保险研究中,标签消耗减少超90%,同时保持决策准确率。
原文摘要 · Abstract (English)
Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes. Recent advances in large language models (LLMs) and other AI systems offer inexpensive auxiliary data, but introduce a new challenge: AI outputs are not direct observations of the target outcomes, but could involve high-dimensional representations with complex and unknown relationships to human labels. Conventional methods leverage AI predictions as direct proxies for true labels, which can be inefficient or unreliable when this relationship is weak or misspecified. We propose Generative Augmented Inference (GAI), a general framework that incorporates AI-generated outputs as informative features for estimating models of human-labeled outcomes. GAI uses an orthogonal moment construction that enables consistent estimation and valid inference with a flexible, nonparametric relationship between LLM-generated outputs and human labels. We establish asymptotic normality and a key dominance result: under random labeling, GAI is optimal within a unified class of debiased estimators-including human-data-only estimators and state-of-the-art debiasing methods-and delivers strict improvements under a mild informativeness condition. Even when the labeled sample is not representative of the target population, an extended variant of GAI still dominates the weighted human-data-only estimator. Empirically, GAI outperforms benchmarks across diverse marketing research settings. In a conjoint analysis, it halves estimation error and reduces human labeling requirements by over 75%. In a pricing study, it consistently outperforms alternative estimators when all methods receive identical auxiliary inputs. In a health insurance study, it saves over 90% of labels while preserving decision accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。