用已有用户数据生成高精度数字分身,提升市场研究效率
Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?

- 从真实面板数据构建个体级大模型分身,覆盖多维度信息深度
- 信息深度达75%熵值时性价比最优,准确率最高达78.8%
- 使用对话历史嵌入可显著提升预测效果,适合营销研究场景
基于大模型的数字分身有望加速市场研究,但现有方法或仅依赖少量人口统计信息,或需专门收集问卷与访谈。本文利用德国社会经济面板(SOEP)数据,构建个体级数字分身,评估了3种开放权重LLM、5种按归一化香农熵排序的信息深度、2种嵌入方式及2种推理模式的组合表现,在500名参与者和183个保留问题上共测试超210万条回答。结果表明,分身质量随信息深度提升,但超过75%熵值后收益递减,该点为成本效益最优的帕累托前沿。将嵌入方式从叙事摘要改为原始对话历史,在所有模型-推理组合中均提升保留准确率;启用显式思维模式则提高排序相关性但不改变准确率。最佳模型在保留集上准确率达78.8%,Fisher-z相关系数r=0.590。研究显示,数字分身应用已不再受制于数据设计,而受限于题项数量、模型选择及少数关键构建决策。
原文摘要 · Abstract (English)
LLM-based digital twins promise to scale and accelerate market research, but most published twins are either coarse persona bots conditioned on a few demographic questions or detailed individual-level twins built on purpose-collected surveys and interview transcripts. Neither setup speaks to the operationally most relevant case for marketing practice: building detailed individual twins from the pre-existing heterogeneous panel data that firms already accumulate through CRM systems, loyalty programs, and repeat surveys. We construct detailed individual-level twins from the German Socio-Economic Panel (SOEP) and evaluate them across a $3 \times 5 \times 2 \times 2$ construction-method grid that covers three open-weights LLMs, five cumulative information depths ranked by normalized Shannon entropy, two embedding methods, and two reasoning modes, scoring over 2.1 million twin responses on 500 participants and 183 held-out questions. Twin quality rises with information depth but with diminishing returns past the 75 percent entropy quartile, which acts as a cost-efficient Pareto point relative to the best-performing 100 percent cells. Switching the embedding from a narrative persona summary to a raw dialog history of past responses raises hold-out accuracy in every model-by-reasoning cell at the 100 percent depth, while an explicit thinking mode raises rank-order correlation without moving accuracy. Best-cell accuracy reaches 78.8 percent and Fisher-$z$ correlation reaches $r = 0.590$ on the SOEP held-out evaluation set. The findings suggest that twin-based market research is no longer gated by data design, but by item volume, model selection, and a small set of construction-level decisions that this paper now maps.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。