用AI生成的韩国虚拟用户模拟真实调查数据,发现效果有限但可诊断。
Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey
- 用大模型生成8000名韩国虚拟用户,对比真实调查数据验证其表现。
- 平均误差15-19个百分点,部分群体差距超52个百分点。
- 校准后仅在极少数真实数据时有优势,不适合直接替代真实调查。
基于大语言模型的合成人格日益被提议作为人类调查受访者的替代方案,但在非英语语境下的系统性验证仍匮乏。本研究利用二次数据,评估了韩国合成人格面板(NVIDIA Nemotron-Personas-Korea)在条件化为Gemini 3.5 Flash(主模型)和EXAONE(对比模型)时,对韩国媒体面板调查(Korea Media Panel Survey)中数字与人工智能服务使用分布的还原能力。每模型生成约8000名分性别与年龄层的虚拟用户,回答原调查中的八项服务使用指标及八项创新性与接受度构念,并与加权调查估计值进行比较。总体平均绝对误差(MAE;RQ1)为15-19个百分点,二元项目均值相关系数为0.69-0.90。按五个社会人口学轴线分析的分段误差(RQ2)同样为15-19个百分点,组间差距最高达52.4/36.2个百分点(Gemini/EXAONE)。误差呈现模型特异性模式:Gemini存在年龄刻板印象且锚定不足,而EXAONE表现出一致的应答倾向偏差。参考年份分析表明,大多数生成式AI高估源于时间错配,而简短形式低估则受表述方式影响。对30%真实数据进行保留校准(RQ3)使性别-年龄单元的MAE大致减半(18.9→8.6,15.9→6.7个百分点),但相同真实子样本的直接估计更精准(3.6个百分点),且校准效果无法跨时间迁移。校准后的面板仅在真实数据极度稀缺(约100个响应)或针对未观测群体时具有优势。叙事条件化优于仅基于人口统计的条件化,但均不及简单的真实数据基线。因此,合成面板并非调查替代品,其价值在于诊断性分析,实际应用仅限于缺乏真实数据的场景。
原文摘要 · Abstract (English)
Synthetic personas based on large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet systematic validation outside English-speaking contexts remains scarce. This secondary-data study evaluates how well a Korean synthetic persona panel (NVIDIA Nemotron-Personas-Korea), conditioned into Gemini 3.5 Flash (primary) and EXAONE (comparison), reproduces digital and AI service-use distributions from the KISDI Korea Media Panel Survey. Sex-and-age-stratified panels of about 8,000 personas per model answered the survey's own items - eight service-use indicators and eight innovativeness and acceptance constructs - and were compared against weighted survey estimates. The overall mean absolute error (MAE; RQ1) was 15-19 percentage points (pp), with binary item-mean correlations of 0.69-0.90. Segment error (RQ2) across five demographic axes was 15-19 pp, with between-group gaps up to 52.4/36.2 pp (Gemini/EXAONE). Errors followed model-specific signatures: an age stereotype with low anchoring (Gemini) versus an acquiescence-consistent level bias (EXAONE). Reference-year analysis was consistent with temporal misalignment driving most generative-AI overestimation, whereas short-form underestimation was framing-sensitive. Holdout calibration on 30% of the real data (RQ3) roughly halved sex-by-age cell MAE (18.9->8.6, 15.9->6.7 pp) - yet direct estimation from the same real subsample was far more accurate (3.6 pp), and the correction did not transfer across time. The calibrated panel retained an advantage only under extremely scarce real data (about 100 responses) and, for one model, for unobserved segments. Persona-narrative conditioning beat demographic-only conditioning, but neither surpassed simple real-data baselines. Synthetic panels are thus not survey substitutes; their value is diagnostic, with operational use confined to settings lacking real data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。